Advances in the Neural Network Quantization: A Comprehensive Review

N/ACitations
Citations of this article
78Readers
Mendeley users who have this article in their library.

Abstract

Artificial intelligence technologies based on deep convolutional neural networks and large language models have made significant breakthroughs in many tasks, such as image recognition, target detection, semantic segmentation, and natural language processing, but also face a conflict between the high computational capacity of the algorithms and limited deployment resources. Quantization, which converts floating-point neural networks into low-bit-width integer networks, is an important and essential technique for efficient deployment and cost reduction in edge computing. This paper analyzes various existing quantization methods, showcases the deployment accuracy of advanced techniques, and discusses the future challenges and trends in this domain.

Cite

CITATION STYLE

APA

Wei, L., Ma, Z., Yang, C., & Yao, Q. (2024, September 1). Advances in the Neural Network Quantization: A Comprehensive Review. Applied Sciences (Switzerland). Multidisciplinary Digital Publishing Institute (MDPI). https://doi.org/10.3390/app14177445

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free