A Survey of Large Language Models for Arabic Language and its Dialects

  • Mashabi M
  • Al-Khalifa S
  • Al-Khalifa H
N/ACitations
Citations of this article
62Readers
Mendeley users who have this article in their library.
Get full text

Abstract

This survey presents a comprehensive review of Large Language Models (LLMs) developed for the Arabic language and its dialects. It categorizes models by architecture (encoder-only, decoder-only, and encoder-decoder) and by linguistic form, including Classical Arabic, Modern Standard Arabic, and Dialectal Arabic. We analyze monolingual, bilingual, and multilingual models, evaluating their performance on tasks such as sentiment analysis, named entity recognition, and question answering. The survey also assesses model openness, considering factors like access to source code, training data, weights, and documentation. Our findings highlight a concentration of resources on MSA, a lack of diverse dialectal datasets, and limited transparency across many models. This work offers the first systematic comparison of openness and linguistic coverage in Arabic LLMs and outlines key challenges and research opportunities to support more inclusive, reproducible, and representative Arabic NLP.

Cite

CITATION STYLE

APA

Mashabi, M., Al-Khalifa, S., & Al-Khalifa, H. (2026). A Survey of Large Language Models for Arabic Language and its Dialects. ACM Transactions on Asian and Low-Resource Language Information Processing. https://doi.org/10.1145/3807946

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free