Abstract
Vanilla Deep Neural Networks (DNN) after training are represented with native floating-point 32 (fp32) weights. We observe that the bit-level sparsity of these weights is very abundant in the mantissa and can be directly exploited to speed up model inference. In this paper, we propose Mortar, an off-line/on-line collaborated approach for fp32 DNN acceleration, which includes two parts: first, an off-line bit sparsification algorithm to construct the target formulation by "mantissa morphing", which maintains higher model accuracy while increasing bit-level sparsity; second, the associating hardware accelerator architecture to speed up the on-line fp32 inference through manipulating the enlarged bit sparsity. We highlight the following results by evaluating various deep learning tasks, including image classification, object detection, video understanding, video & image super-resolution, etc.: We (1) increase bit-level sparsity up to 1.28∼2.51x with only a negligible-0.09∼0.23% accuracy loss, (2) maintain on average 3.55% higher model accuracy while increasing more bit-level sparsity than the baseline, (3)and our hardware accelerator outperforms up to 4.8x over the baseline, with an area of 0.031 mm2and power of 68.58 mW.
Author supplied keywords
Cite
CITATION STYLE
Gao, Y., Li, H., Zhang, K., Yu, X., & Lu, H. (2023). Mortar: Morphing the Bit Level Sparsity for General Purpose Deep Learning Acceleration. In Proceedings of the Asia and South Pacific Design Automation Conference, ASP-DAC (pp. 739–744). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1145/3566097.3567868
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.