Abstract
This review provides a concise overview of key transformer-based language models, including bidirectional encoder representations from transformers (BERT), generative pre-trained transformer 3 (GPT-3), robustly optimized BERT pretraining approach (RoBERTa), a lite BERT (ALBERT), text-to-text transfer transformer (T5), generative pre-trained transformer 4 (GPT-4), and extra large neural network (XLNet). These models have significantly advanced natural language processing (NLP) capabilities, each bringing unique contributions to the field. We delve into BERT’s bidirectional context understanding, GPT-3’s versatility with 175 billion parameters, and RoBERTa’s optimization of BERT. ALBERT emphasizes model efficiency, T5 introduces a text-to-text framework, and GPT-4, with 170 trillion parameters, excels in multimodal tasks. Safety considerations are highlighted, especially in GPT-4. Additionally, XLNet’s permutation-based training achieves bidirectional context understanding. The motivations, advancements, and challenges of these models are explored, offering insights into the evolving landscape of large-scale language models.
Author supplied keywords
Cite
CITATION STYLE
Briouya, A., Briouya, H., & Choukri, A. (2024). Overview of the progression of state-of-the-art language models. Telkomnika (Telecommunication Computing Electronics and Control), 22(4), 897–909. https://doi.org/10.12928/TELKOMNIKA.v22i4.25936
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.