scieee AI-readable full text Open interactive document viewer

QeiHaN: An energy-efficient DNN accelerator that leverages log quantization in NDP architectures

Khabbazan, Bahareh,Riera Villanueva, Marc,González Colás, Antonio María

Abstract

The constant growth of DNNs makes them challenging to implement and run efficiently on traditional computecentric architectures. Some works have attempted to enhance accelerators by adding more compute units and on-chip buffers, but they often worsen the memory issue due to increased bandwidth demands. Memory-centric designs based on Near-Data Processing (NDP) have been proposed to mitigate this problem by moving computations closer to the memory hierarchy. Leveraging 3D-stacked memory for its storage density and near-memory processing capabilities, this paper introduces QeiHaN, a hardware accelerator that optimizes DNN inference efficiency. QeiHaN employs a 3D-stacked memory-centric weight storage scheme combined with a logarithmic quantization of activations, resulting in reduced memory accesses by 25%. Evaluation demonstrates significant speedup and energy savings compared to a Neurocube-like accelerator across various DNNs.

Full text

QeiHaN: An Energy-Efficient DNN Accelerator that Leverages Logarithmic Quantization in Near-Data Processing Architectures Bahareh Khabbazan, Marc Riera, Antonio González {bahareh.khabbazan,marc.riera.villanueva,antonio.gonzalez}@upc.edu Universitat Politècnica de Catalunya (UPC) Barcelona, Spain ABSTRACT The constant growth of DNNs makes them challenging to implement and run efficiently on traditional compute-centric architectures. Some works have attempted to enhance accelerators by adding more compute units and on-chip buffers, but they often worsen the memory issue due to increased bandwidth demands. Memory-centric designs based on Near-Data Processing (NDP) have been proposed to mitigate this problem by moving computations closer to the memory hierarchy. Leveraging 3D-stacked memory for its storage density and near-memory processing capabilities, this paper introduces QeiHaN, a hardware accelerator that optimizes DNN inference efficiency. QeiHaN employs a 3D-stacked memorycentric weight storage scheme combined with a logarithmic quantization of activations, resulting in reduced memory accesses by 25%. Evaluation demonstrates significant speedup and energy savings compared to a Neurocube-like accelerator across various DNNs. KEYWORDS DNN, NDP, Accelerators, Quantization, Exponential, Transformer 1 INTRODUCTION DNNs are a powerful solution to a broad range of machine learning applications at the expense of high computational cost, memory requirements, and energy consumption. The constant growth of DNNs makes it difficult to efficiently implement them even on modern accelerators. The concept of NDP involves moving computations closer to the memory to address the memory wall problem. Traditional DNN accelerators tend to focus their area on the processing elements (PEs) that are responsible to speed-up the frequent dot-product operations. However, they often suffer from memory bandwidth limitations and energy inefficiency due to data movements. Data transfers consume a significant portion of energy in such systems. Recent research suggests that many data movements are caused by simple functions that could be implemented in hardware. This motivates a transition from compute-centric to data-centric architectures for data-intensive applications. NDP architectures based on 3D-stacked memory, such as Hybrid Memory Cube (HMC), High Bandwidth Memory (HBM), and Wide I/O, aim to break the memory wall by increasing storage capacity, memory bandwidth, and reducing power consumption. These architectures promise benefits for scalable DNN accelerators due to their high bandwidth and parallel memory access. Examples of NDP architectures like Neurocube [ 2 ] and TETRIS [ 1 ] offer promising performance and energy efficiency for DNN acceleration, but there 0% 5% 10% 15% 20% 25% 30% -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 LogQuant Activations Exponents AlexNet PTBLM Transformer BERT-Base BERT-Large Figure 1: Histograms of the LOG2 Quantization (LogQuant) of activations from all FC and CON layers of a set of DNNs. is room for improvement. Challenges include redesigning on-chip buffers, optimizing dataflow scheduling and partitioning of DNN computations, and dealing with thermal constraints. This work shows how to exploit a logarithmic base-2 quantization (LOG2) of activations on FC and CONV layers of DNNs. We propose an implicit in-memory bit-shifting of DNN weights to reduce memory movements. Weights are uniformly quantized and stored at the bit-level granularity into different memory banks to exploit the parallelism of 3D-stacked architectures. Next, we propose a mechanism to skip accessing the bits of the weights that are not useful due to the right bit-shifting of the negative exponents of the LOG2 quantized activations. Then, we present QeiHaN, an NDP accelerator that implements our LOG2 quantization-shifting engine and efficient weight storage scheme for high-performance low-energy DNN inference. QeiHaN is designed on top of a Neurocube-like architecture with an enhanced input stationary dataflow, introducing modest hardware modifications. QeiHaN requires a small set of comparators and integer adders to perform the LOG2 quantization. Then, we also replace the multipliers by simple bit-shift logic reducing the computational cost. Our experimental results show that the overheads are minimal compared to the savings in memory accesses and multiplications. 2 LOG2 QUANTIZATION ANALYSIS Linear quantization suffers accuracy loss due to non-uniform activation and weight distributions, particularly in complex DNNs with multiple layers. LOG2 quantization better accommodates these distributions, especially for activations, by exploiting their characteristics. The main benefit of the LOG2 quantization is that it not only reduces the numerical precision but also eliminates the bulky digital multipliers by using simple shift and ADD operations. A quantitative analysis of quantized activation exponents reveals huge percentages of negative values. QeiHaN exploits these observations yielding memory savings and optimized storage. Figure 1 shows that across DNNs over 71% of activations exhibit negative 1 © 2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes,creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. https://dx.doi.org/10.1109/PACT58117.2023.00036 Khabbazan et al. PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC PE RVC PE RVC PE RVC PECPEC Logic Die PE I/O Buffer LOG2-Quant Decoder & Shifter OB R RPEC SFU Weights Buffer IB ADD0... ADDd-1 ADDd ADD0... ADDd-1 ADDd Figure 2: Architecture of a Processing Element (PE) of QeiHaN. exponents. Note that the positive exponents will lead to a shift to the left, while negative exponents result in a shift to the right discarding the least significant bits (LSB) of the weights that are multiplied by the corresponding activation. 3 QeiHaN ACCELERATOR The goal of QeiHaN is to optimize memory pressure by exploiting the LOG2 quantization of activations and an efficient weight storage scheme. QeiHaN is based on NDP architectures that leverage 3D stacked memory for efficient DNN inference. Figure 2 outlines the QeiHaN accelerator. Each logic die tile includes a PE, a Vault Controller (VC), a Router (R), and a PE Controller (PEC). The PE is the core responsible for DNN acceleration. The key PE components are SRAM blocks for storing inputs (IB), outputs (OB) and weights (WB), the LOG2 Quantization (LOG2-Quant) unit, Weight Decoder and Shifter (D&S) unit, ADD array, and Special Function Unit (SFU). First, in the (D&S) unit, weights that multiply non-zero activations are decoded from a compressed stream and bit-shifted by appending the necessary amount of zeros based on the exponent from the LOG2-Quant unit. In order to use this unit efficiently, QeiHaN reorganizes the weights in-memory to a bit-level granularity. Then, the ADD Array is made of dindependent ADD units that are used to accumulate the products of each activation by the corresponding weights. According to the sign of the activation value, not the exponent, the bit-shifted weight is added/subtracted to/from the partial outputs computed in previous cycles and stored in the OB. LOG2 quantization obviates multipliers; thus, partial outputs load from OB, and bit-shifted weights originate from (D&S). In a single execution all adders compute partial outputs related to the same activation from ddifferent convolutional kernels or output neurons. Finally, the SFU is composed of units to perform non-linear activation functions, pooling, and normalization, among others. Regarding the memory organization in QeiHaN, input activations are divided by channel across vaults, while vaults allocate segments of partial outputs for all channels. To manage the high dimensionality of inputs/outputs, a blocking approach is employed, breaking IFM and OFM into smaller blocks per channel. Our layout to store weights in DRAM involves interleaving the bits across banks and partitions within vaults. This arrangement simplifies the implicit bit-shifting scheme. QeiHaN’s data remapping optimizes bank-level parallelism in 3D-stacked DRAM-based operations, enhancing bandwidth by overlapping requests to different banks. QeiHaN uses an enhanced input stationary (IS) dataflow coupled with a blocking scheme to efficiently exploit the LOG2 quantization of the input activations, minimizing the memory accesses to both Start Layer End Layer More Blocks? Read IB from DRAM Die Write IB in I/O Buffer Fit IB? Read Input from I/O Buffer QLog2 & Clip Input Zero Input? More Inputs? Read Bits of Weights from DRAM Decode & Shift Weights Read Outputs from I/O Buffer Accumulate Results Write Outputs to I/O Buffer Reduction Tree of Outputs Activation Pooling More Weights? Write OBs to DRAM Yes No Yes Yes Yes No No No No Yes De-quantization Figure 3: Dataflow and Execution scheme. weights and activations. Figure 3 illustrates the dataflow of the QeiHaN accelerator with a flowchart. The proposed dataflow includes three main stages marked in different colors: Preprocessing (gray), execution (orange), and postprocessing (blue). 4 EVALUATION We evaluate the area and energy consumption of the accelerator by implementing and synthesizing the logic components using Design Compiler and the technology library of 28/32nm from Synopsys. Also, the energy consumption of the 3D-stacked memory is estimated by using an HMC configuration of DRAMSim3. For a fair comparison, we set most of the configuration parameters to match the Neurocube [ 2 ] baseline. On average for our set of DNNs, QeiHaN reduces the total DRAM accesses by 72.4% over Neurocube baseline. Compared to Neurocube, QeiHaN provides consistent speedups for the five DNNs achieving an average performance improvement of 4.25xand energy savings of 3.52x. Regarding the area, the area overhead of QeiHaN in the logic die due to 16 PEs is 0.389mm 2 . In comparison, Neurocube extra storage and multipliers result in 20% more area than QeiHaN at the same technology node. 5 CONCLUSIONS This paper reveals that modern DNNs exhibit notable negative exponent distribution of activations post logarithmic quantization, resulting in increased right bit-shift operations. To exploit this observation we introduce QeiHaN, a new 3D-stacked DRAM-based NDP accelerator. QeiHaN replaces multiplications and reduces memory accesses to only the useful bits of the weights, by implementing an implicit in-memory bit-shifting. Our experimental results show that QeiHaN provides 3.5xenergy savings and 4.3xspeedup with negligible accuracy loss and lower area than Neurocube. 6 ACKNOWLEDGEMENT This work has been supported by the CoCoUnit ERC Advanced Grant of the EU’s Horizon 2020 program (grant No 833057), the Spanish State Research Agency (MCIN/AEI) under grant PID2020113172RB-I00, and the ICREA Academia program. REFERENCES [1] Mingyu Gao, Jing Pu, Xuan Yang, Mark Horowitz, and Christos Kozyrakis. 2017. Tetris: Scalable and efficient neural network acceleration with 3d memory. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems. 751–764. [2] Duckhwan Kim, Jaeha Kung, Sek Chai, Sudhakar Yalamanchili, and Saibal Mukhopadhyay. 2016. Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory. ACM SIGARCH Computer Architecture News 44, 3 (2016), 380–392. 2