Evaluating Arabic Sentiment Analysis With GPT-4o: A Comparative Study of Raw and Preprocessed Text

1Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Generative Pretrained Transformer 4 Omni (GPT-4o) is a large language model that has contributed to natural language processing by modeling complex linguistic patterns. However, given the scarcity of labeled Arabic data and the growing interest in deploying large language models without task-specific fine-tuning, there is a lack of studies that evaluate the effectiveness and reliability of GPT-4o in processing Arabic texts in the domain of sentiment analysis under zero-shot settings. To address this gap, experiments on original and preprocessed texts using two Arabic datasets—restaurant reviews and multi-topic tweets—were conducted. The zero-shot performance of GPT-4o was compared with Arabic Bidirectional Encoder Representations from Transformers, Deep Bidirectional Transformers for Arabic, and Multilingual Bidirectional Encoder Representations from Transformers. Overall, GPT-4o outperformed other models, with accuracies of 85% for preprocessed restaurant reviews and 77% for original tweets. The tweet dataset was more challenging, indicating GPT-4o’s performance is affected by informal and noisy social media content. These findings highlight important considerations for adoption of GPT-4o in Arabic sentiment analysis applications.

Cite

CITATION STYLE

APA

Wali, A., Alahmadi, A., Aljohani, A., & Alharbi, A. (2026). Evaluating Arabic Sentiment Analysis With GPT-4o: A Comparative Study of Raw and Preprocessed Text. Journal of Cases on Information Technology, 28(1). https://doi.org/10.4018/JCIT.402749

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free