Crossbar-Aligned & Integer-Only Neural Network Compression for Efficient in-Memory Acceleration

5Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Crossbar-based In-Memory Computing (IMC) accelerators preload the entire Deep Neural Network (DNN) into crossbars before inference. However, devices with limited crossbars cannot infer increasingly complex models. IMC-pruning can reduce the usage of crossbars, but current methods need expensive extra hardware for data alignment. Meanwhile, quantization can represent weights of DNNs by integers, but they employ non-integer scaling factors to ensure accuracy, requiring costly multipliers. In this paper, we first propose crossbar-aligned pruning to reduce the usage of crossbars without hardware overhead. Then, we introduce a quantization scheme to avoid multipliers in IMC devices. Finally, we design a learning method to complete above two schemes and cultivate an optimal compact DNN with high accuracy and large sparsity during training. Experiments demonstrate that our framework, compared to state-of-the-art methods, achieves larger sparsity and lower power consumption with higher accuracy. We even improve the accuracy by 0.43% for VGG-16 with an 88.25% sparsity rate on the Cifar-10 dataset. Compared to the original model, we reduce computing power and area by 19.8x and 18.8x, respectively.

Cite

CITATION STYLE

APA

Huai, S., Liu, D., Luo, X., Chen, H., Liu, W., & Subramaniam, R. (2023). Crossbar-Aligned & Integer-Only Neural Network Compression for Efficient in-Memory Acceleration. In Proceedings of the Asia and South Pacific Design Automation Conference, ASP-DAC (pp. 234–239). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1145/3566097.3567856

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free