Evaluating Large Language Models on Sentiment Analysis in Arabic Dialects

2Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Despite recent progress in large language models (LLMs), their performance on Arabic dialects remains underexplored, particularly in the context of sentiment analysis. This study presents a comparative evaluation of three LLMs, DeepSeek-R1, Qwen2.5, and LLaMA-3, on sentiment classification across Modern Standard Arabic (MSA), Saudi dialect and Darija. We construct a balanced sentiment dataset by translating and validating MSA hotel reviews into Saudi dialect and Darija. Using parameter-efficient fine-tuning (LoRA) and dialect-specific prompts, we assess each model under matched and mismatched prompting conditions. Experimental results show that Qwen2.5 achieves the highest macro F1 score of 79% on Darija input using MSA prompts, while DeepSeek performs best when prompted in the input dialect, reaching 71% on Saudi dialect. LLaMA-3 exhibits stable performance across prompt variations, with 75% macro F1 on Darija input under MSA prompting. Dialect-aware prompting consistently improves classification accuracy, particularly for neutral and negative sentiment classes.

Cite

CITATION STYLE

APA

Alharbi, M., Ezzini, S., Ranasinghe, T., Hettiarachchi, H., & Mitkov, R. (2025). Evaluating Large Language Models on Sentiment Analysis in Arabic Dialects. In International Conference Recent Advances in Natural Language Processing, RANLP (pp. 67–74). Incoma Ltd. https://doi.org/10.26615/978-954-452-098-4-008

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free