Efficient LLMs Training and Inference: An Introduction

22Citations
Citations of this article
28Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

ChatGPT was released in late November 2022, making a significant impact globally. Following this release, numerous domestic and international open-source projects for large model training emerged, including Alpaca, BOOLM, LLaMA, ChatGLM, DeepSpeedChat, and ColossalChat. Both academia and industry have a growing need to train large models for optimizing downstream tasks. Research has demonstrated that instruction tuning using high-quality training data can enable models with over 10 billion parameters to exhibit emergent abilities, particularly in complex reasoning tasks. Training a model with tens of billions of parameters is often necessary to improve business metrics using large model technology. However, most research teams lack the extensive computing resources required for such large-scale models, often possessing only a few V100 (32G) GPUs and occasionally a few A100s. Consequently, training or inferring a large model of this magnitude is impractical under limited computational resources. Therefore, adopting optimization strategies during the training and inference stages is essential to address these challenges. This survey summarizes a series of optimization techniques for large models during the training and inference phases, enabling the training of large models even with constrained computational resources.

Cite

CITATION STYLE

APA

Li, R., Fu, D., Shi, C., Huang, Z., & Lu, G. (2025). Efficient LLMs Training and Inference: An Introduction. IEEE Access, 13, 32944–32970. https://doi.org/10.1109/ACCESS.2024.3501358

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free