Abstract
This paper addresses the multilingual language understanding of ChatGPT-3.5 and 4 to investigate their performance with respect to languages with different degrees of prevalence on the internet. ChatGPT’s training data mostly consists of website content. As the language distribution is unevenly allocated and a low number of languages is used on websites this should impact performance. Both ChatGPT versions should rate reviews between 1 to 5 stars based solely on the product description and the review texts. Therefore, 500 e-commerce reviews are collected for each of five languages: English, German, Dutch, Korean and Hindi, which are evenly distributed at 100 reviews per star rating. The evaluation methods and metrics used in this study include t-tests, confusion matrices, macro F1 values and a defined cumulative star deviation. The results indicate a significant correlation between the degree of dissemination and the accuracy of the ChatGPT-3.5 evaluation. In direct comparison, ChatGPT-4 shows superior accuracy in all languages studied, while maintaining acceptable performance in less represented languages. The hypothesis that ChatGPT-4 scoring accuracy increases with an increase in the number of words in reviews in less represented languages could not be confirmed. These findings illustrate the influence of the selected language on the interaction with ChatGPT and its language comprehension, which suggests that multilingualism should be given greater consideration in the future development and optimization of large language models.
Author supplied keywords
Cite
CITATION STYLE
Erös, B., Gritsch, C., Tick, A., & Rosenberger, P. (2024). Comparison of effectiveness between ChatGPT 3.5 and 4 in understanding different natural languages. Journal of Intelligence Studies in Business, 14(2), 77–97. https://doi.org/10.37380/jisib.v14.i2.2547
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.