Overview of the progression of state-of-the-art language models

N/ACitations
Citations of this article
27Readers
Mendeley users who have this article in their library.

Abstract

This review provides a concise overview of key transformer-based language models, including bidirectional encoder representations from transformers (BERT), generative pre-trained transformer 3 (GPT-3), robustly optimized BERT pretraining approach (RoBERTa), a lite BERT (ALBERT), text-to-text transfer transformer (T5), generative pre-trained transformer 4 (GPT-4), and extra large neural network (XLNet). These models have significantly advanced natural language processing (NLP) capabilities, each bringing unique contributions to the field. We delve into BERT’s bidirectional context understanding, GPT-3’s versatility with 175 billion parameters, and RoBERTa’s optimization of BERT. ALBERT emphasizes model efficiency, T5 introduces a text-to-text framework, and GPT-4, with 170 trillion parameters, excels in multimodal tasks. Safety considerations are highlighted, especially in GPT-4. Additionally, XLNet’s permutation-based training achieves bidirectional context understanding. The motivations, advancements, and challenges of these models are explored, offering insights into the evolving landscape of large-scale language models.

Cite

CITATION STYLE

APA

Briouya, A., Briouya, H., & Choukri, A. (2024). Overview of the progression of state-of-the-art language models. Telkomnika (Telecommunication Computing Electronics and Control), 22(4), 897–909. https://doi.org/10.12928/TELKOMNIKA.v22i4.25936

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free