An Empirical Study of Speculative Decoding for Small Language Models

0Citations
Citations of this article
3Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Speculative decoding has emerged as a promising approach to accelerate Large Language Model inference. However, existing research has predominantly focused on 7B-70B parameters models, leaving a critical knowledge gap for small language models (1-2B parameters) that are increasingly important for edge computing and agentic AI systems. This paper presents the first comprehensive empirical study of speculative decoding techniques for small language models. We evaluate five distinct method categories across three representative model families and reveal that drafting overhead, rather than draft quality, becomes the primary bottleneck fundamentally limiting acceleration of small models. We demonstrate that traditional independent drafting fails completely due to the suboptimal architecture of available drafters, while self-drafting methods achieve meaningful acceleration only when employing sufficiently efficient draft modules. In contrast, retrieval-based methods with negligible computational overhead yield consistent gains. Based on these insights, we establish practical guidelines for effective small model acceleration.

Cite

CITATION STYLE

APA

Mainardi, L., Sandikci, S., & Vanschoren, J. (2026). An Empirical Study of Speculative Decoding for Small Language Models. In EACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference, Vol. 1 - (Long Papers) (Vol. 1, pp. 5483–5497). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2026.eacl-long.255

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free