Benchmarking Large Language Models from Open and Closed Source Models to Apply Data Annotation for Free-Text Criteria in Healthcare

5Citations
Citations of this article
15Readers
Mendeley users who have this article in their library.

Abstract

Large language models (LLMs) hold the potential to significantly enhance data annotation for free-text healthcare records. However, ensuring their accuracy and reliability is critical, especially in clinical research applications requiring the extraction of patient characteristics. This study introduces a novel evaluation framework based on Multi-Criteria Decision Analysis (MCDA) and the Order of Preference by Similarity to Ideal Solution (TOPSIS) technique, designed to benchmark LLMs on their annotation quality. The framework defines ten evaluation metrics across key criteria such as age, gender, BMI, disease presence, and blood markers (e.g., white blood count and platelets). Using this methodology, we assessed leading open source and commercial LLMs, achieving accuracy scores of 0.59, 1, 0.84, 0.56, and 0.92, respectively, for the specified criteria. Our work not only provides a rigorous framework for evaluating LLM capabilities in healthcare data annotation but also highlights their current performance limitations and strengths. By offering a comprehensive benchmarking approach, we aim to support responsible adoption and decision-making in healthcare applications.

Cite

CITATION STYLE

APA

Nemati, A., Assadi Shalmani, M., Lu, Q., & Luo, J. (2025). Benchmarking Large Language Models from Open and Closed Source Models to Apply Data Annotation for Free-Text Criteria in Healthcare. Future Internet, 17(4). https://doi.org/10.3390/fi17040138

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free