Abstract
Mixed-Precision Quantization (MPQ) has become a key technique for deploying deep neural networks on resource-constrained IoT edge devices, enabling efficient TinyML applications while maintaining accuracy. However, current MPQ solutions face the following major limitations: insufficient flexibility in supporting diverse neural architectures and lack of energy-efficient hardware acceleration tailored for edge deployment. To address these challenges, we propose an integrated software-hardware co-design framework that optimizes Mixed-Precision Quantized Neural Networks for edge AI inference. Our framework introduces following key innovations: an adaptive reorderd pack scheme to improve data utilization, dynamic data flow schemes for low-bit-width computationally intensive operators, and customized SIMD (Single-Instruction, Multiple-Data) instruction extensions for RISC-V processors optimized for flexible mixed-precision computations. Experimental results demonstrate that our framework achieves 1.23×–1.58× speedup for 2–8 bit operations compared to existing RISC-V implementations while maintaining excellent energy efficiency of 1,204 GOPS/W. Our design also keeps the power range to 2.74 mW, making it particularly suitable for cost-sensitive edge deployments. These advances provide a practical solution for next-generation IoT systems requiring adaptive precision and ultra-low-power execution on edge devices.
Author supplied keywords
Cite
CITATION STYLE
Lu, T., Ding, J., Zhang, H., Jiang, B., Xu, W., & Chai, Z. (2025). Efficient Flexible Edge Inference for Mixed-Precision Quantized DNN using Customized RISC-V Core. ACM Transactions on Architecture and Code Optimization, 22(4). https://doi.org/10.1145/3768630
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.