Abstract
Despite recent progress in large language models (LLMs), their performance on Arabic dialects remains underexplored, particularly in the context of sentiment analysis. This study presents a comparative evaluation of three LLMs, DeepSeek-R1, Qwen2.5, and LLaMA-3, on sentiment classification across Modern Standard Arabic (MSA), Saudi dialect and Darija. We construct a balanced sentiment dataset by translating and validating MSA hotel reviews into Saudi dialect and Darija. Using parameter-efficient fine-tuning (LoRA) and dialect-specific prompts, we assess each model under matched and mismatched prompting conditions. Experimental results show that Qwen2.5 achieves the highest macro F1 score of 79% on Darija input using MSA prompts, while DeepSeek performs best when prompted in the input dialect, reaching 71% on Saudi dialect. LLaMA-3 exhibits stable performance across prompt variations, with 75% macro F1 on Darija input under MSA prompting. Dialect-aware prompting consistently improves classification accuracy, particularly for neutral and negative sentiment classes.
Cite
CITATION STYLE
Alharbi, M., Ezzini, S., Ranasinghe, T., Hettiarachchi, H., & Mitkov, R. (2025). Evaluating Large Language Models on Sentiment Analysis in Arabic Dialects. In International Conference Recent Advances in Natural Language Processing, RANLP (pp. 67–74). Incoma Ltd. https://doi.org/10.26615/978-954-452-098-4-008
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.