Mortar: Morphing the Bit Level Sparsity for General Purpose Deep Learning Acceleration

4Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Vanilla Deep Neural Networks (DNN) after training are represented with native floating-point 32 (fp32) weights. We observe that the bit-level sparsity of these weights is very abundant in the mantissa and can be directly exploited to speed up model inference. In this paper, we propose Mortar, an off-line/on-line collaborated approach for fp32 DNN acceleration, which includes two parts: first, an off-line bit sparsification algorithm to construct the target formulation by "mantissa morphing", which maintains higher model accuracy while increasing bit-level sparsity; second, the associating hardware accelerator architecture to speed up the on-line fp32 inference through manipulating the enlarged bit sparsity. We highlight the following results by evaluating various deep learning tasks, including image classification, object detection, video understanding, video & image super-resolution, etc.: We (1) increase bit-level sparsity up to 1.28∼2.51x with only a negligible-0.09∼0.23% accuracy loss, (2) maintain on average 3.55% higher model accuracy while increasing more bit-level sparsity than the baseline, (3)and our hardware accelerator outperforms up to 4.8x over the baseline, with an area of 0.031 mm2and power of 68.58 mW.

Cite

CITATION STYLE

APA

Gao, Y., Li, H., Zhang, K., Yu, X., & Lu, H. (2023). Mortar: Morphing the Bit Level Sparsity for General Purpose Deep Learning Acceleration. In Proceedings of the Asia and South Pacific Design Automation Conference, ASP-DAC (pp. 739–744). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1145/3566097.3567868

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free