A critical review of large language models: Sensitivity, bias, and the path toward specialized AI

41Citations
Citations of this article
56Readers
Mendeley users who have this article in their library.

Abstract

This paper examines the comparative effectiveness of a specialized compiled language model and a general-purpose model such as OpenAI’s GPT-3.5 in detecting sustainable development goals (SDGs) within text data. It presents a critical review of large language models (LLMs), addressing challenges related to bias and sensitivity. The necessity of specialized training for precise, unbiased analysis is underlined. A case study using a company descriptions data set offers insight into the differences between the GPT-3.5 model and the specialized SDG detection model. While GPT-3.5 boasts broader coverage, it may identify SDGs with limited relevance to the companies’ activities. In contrast, the specialized model zeroes in on highly pertinent SDGs. The importance of thoughtful model selection is emphasized, taking into account task requirements, cost, complexity, and transparency. Despite the versatility of LLMs, the use of specialized models is suggested for tasks demanding precision and accuracy. The study concludes by encouraging further research to find a balance between the capabilities of LLMs and the need for domain-specific expertise and interpretability.

Cite

CITATION STYLE

APA

Hajikhani, A., & Cole, C. (2024). A critical review of large language models: Sensitivity, bias, and the path toward specialized AI. Quantitative Science Studies, 5(3), 736–756. https://doi.org/10.1162/qss_a_00310

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free