Abstract
Large language models are transforming industries but face challenges due to their high computational and energy demands. Model compression via quantization mitigates these barriers by reducing the bit precision of parameters and arithmetic operations, enabling deployment on resource-constrained devices like smartphones and edge platforms. This article focuses on quantization applied to transformer decoders, which are critical for tasks, such as text generation and conversational artificial intelligence. Unlike encoders, decoders are constrained by memory due to their sequential processing nature and low arithmetic intensity. We propose optimizations targeting inference on low-power CPUs, emphasizing efficient linear layers with quantized data/arithmetic and cache optimization. Using two representative ARM and RISC-V platforms, we present optimized mixed-precision implementations of the matrix multiplication that outperform the instance of that computational kernel in popular libraries, such as basic linear algebra subprograms infrastructure software, XNNPACK and ARMCL. This work thus advances the understanding of the impact of quantization on transformer decoder efficiency, energy consumption and precision in edge environments.
Author supplied keywords
Cite
CITATION STYLE
Martínez, H., Catalán, S., Castelló, A., Mestre, J. I., & Quintana-Ortí, E. S. (2025). Latency-Critical Quantized Inference With Transformer Decoders on ARM and RISC-V CPUs. IEEE Internet of Things Journal, 12(13), 25676–25690. https://doi.org/10.1109/JIOT.2025.3560382
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.