Abstract
Large language models predominantly reflect Western cultures, largely due to the dominance of English-centric training data. This imbalance presents a significant challenge, as LLMs are increasingly used across diverse contexts without adequate evaluation of their cultural competence in non-English languages, including Persian. To address this gap, we introduce PERCUL, a carefully constructed dataset designed to assess the sensitivity of LLMs toward Persian culture. PERCUL features story-based, multiple-choice questions that capture culturally nuanced scenarios. Unlike existing benchmarks, PERCUL is curated with input from native Persian annotators to ensure authenticity and to prevent the use of translation as a shortcut. We evaluate several state-of-the-art multilingual and Persian-specific LLMs, establishing a foundation for future research in cross-cultural NLP evaluation. Our experiments demonstrate a 11.3% gap between best closed source model and layperson baseline while the gap increases to 21.3% by using the best open-weight model. You can access the dataset from here: https://huggingface.co/datasets/teias-ai/percul.
Cite
CITATION STYLE
Monazzah, E. M., Rahimzadeh, V., Yaghoobzadeh, Y., Shakery, A., & Pilehvar, M. T. (2025). PERCUL: A Story-Driven Cultural Evaluation of LLMs in Persian. In Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025 (Vol. 1, pp. 12670–12687). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.naacl-long.631
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.