Efficient Inference for Edge Large Language Models: A Survey

  • Cai G
  • Tian R
  • Yang L
  • et al.
N/ACitations
Citations of this article
17Readers
Mendeley users who have this article in their library.

Abstract

Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing. Their massive computational and memory requirements often necessitate cloud-based deployment, introducing challenges related to cost, latency, privacy, and network reliability. Deploying on-device LLMs alleviates these challenges, but is hindered by the severe resource constraints of edge hardware. This survey reviews efficient inference techniques for edge LLMs, with a focus on two key strategies of speculative decoding and model offloading. We categorize strategies into single-device and multi-device types, systematically analyzing the principles, recent advancements, implementations, and support within edge frameworks. Finally, we highlight the open challenges and future research directions that will advance the field of edge LLM inference.

Cite

CITATION STYLE

APA

Cai, G., Tian, R., Yang, L., Jia, Y., Li, L., & Wang, J. (2026). Efficient Inference for Edge Large Language Models: A Survey. Tsinghua Science and Technology, 31(3), 1365–1380. https://doi.org/10.26599/tst.2025.9010166

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free