Abstract
Crossbar-based In-Memory Computing (IMC) accelerators preload the entire Deep Neural Network (DNN) into crossbars before inference. However, devices with limited crossbars cannot infer increasingly complex models. IMC-pruning can reduce the usage of crossbars, but current methods need expensive extra hardware for data alignment. Meanwhile, quantization can represent weights of DNNs by integers, but they employ non-integer scaling factors to ensure accuracy, requiring costly multipliers. In this paper, we first propose crossbar-aligned pruning to reduce the usage of crossbars without hardware overhead. Then, we introduce a quantization scheme to avoid multipliers in IMC devices. Finally, we design a learning method to complete above two schemes and cultivate an optimal compact DNN with high accuracy and large sparsity during training. Experiments demonstrate that our framework, compared to state-of-the-art methods, achieves larger sparsity and lower power consumption with higher accuracy. We even improve the accuracy by 0.43% for VGG-16 with an 88.25% sparsity rate on the Cifar-10 dataset. Compared to the original model, we reduce computing power and area by 19.8x and 18.8x, respectively.
Author supplied keywords
Cite
CITATION STYLE
Huai, S., Liu, D., Luo, X., Chen, H., Liu, W., & Subramaniam, R. (2023). Crossbar-Aligned & Integer-Only Neural Network Compression for Efficient in-Memory Acceleration. In Proceedings of the Asia and South Pacific Design Automation Conference, ASP-DAC (pp. 234–239). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1145/3566097.3567856
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.