Latency-Critical Quantized Inference With Transformer Decoders on ARM and RISC-V CPUs

5Citations
Citations of this article
13Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Large language models are transforming industries but face challenges due to their high computational and energy demands. Model compression via quantization mitigates these barriers by reducing the bit precision of parameters and arithmetic operations, enabling deployment on resource-constrained devices like smartphones and edge platforms. This article focuses on quantization applied to transformer decoders, which are critical for tasks, such as text generation and conversational artificial intelligence. Unlike encoders, decoders are constrained by memory due to their sequential processing nature and low arithmetic intensity. We propose optimizations targeting inference on low-power CPUs, emphasizing efficient linear layers with quantized data/arithmetic and cache optimization. Using two representative ARM and RISC-V platforms, we present optimized mixed-precision implementations of the matrix multiplication that outperform the instance of that computational kernel in popular libraries, such as basic linear algebra subprograms infrastructure software, XNNPACK and ARMCL. This work thus advances the understanding of the impact of quantization on transformer decoder efficiency, energy consumption and precision in edge environments.

Cite

CITATION STYLE

APA

Martínez, H., Catalán, S., Castelló, A., Mestre, J. I., & Quintana-Ortí, E. S. (2025). Latency-Critical Quantized Inference With Transformer Decoders on ARM and RISC-V CPUs. IEEE Internet of Things Journal, 12(13), 25676–25690. https://doi.org/10.1109/JIOT.2025.3560382

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free