Using Contextually Aligned Online Reviews to Measure LLMs’ Performance Disparities Across Language Varieties

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.
Get full text

Abstract

A language can have different varieties. These varieties can affect the performance of natural language processing (NLP) models, including large language models (LLMs), which are often trained on data from widely spoken varieties. This paper introduces a novel and cost-effective approach to benchmark model performance across language varieties. We argue that international online review platforms, such as Booking.com, can serve as effective data sources for constructing datasets that capture comments in different language varieties from similar real-world scenarios, like reviews for the same hotel with the same rating using the same language (e.g., Mandarin Chinese) but different language varieties (e.g., Taiwan Mandarin, Mainland Mandarin). To prove this concept, we constructed a contextually aligned dataset comprising reviews in Taiwan Mandarin and Mainland Mandarin and tested six LLMs in a sentiment analysis task. Our results show that LLMs consistently underperform in Taiwan Mandarin.

Cite

CITATION STYLE

APA

Tang, Z., Huang, C. Y., Li, T. C., Ng, H. Y. S., Huang, H. H., & Huang, T. H. (2025). Using Contextually Aligned Online Reviews to Measure LLMs’ Performance Disparities Across Language Varieties. In Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025 (Vol. 2, pp. 342–355). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.naacl-short.29

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free