Human evaluation of large language models in healthcare: gaps, challenges, and the need for standardization

7Citations
Citations of this article
25Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Publications related to experimentation with Large Language Models (LLMs) in healthcare are rapidly increasing. While human evaluation remains the gold standard for evaluating LLMs, there is still a lack of standardization in its implementation. In this review article, we systematically examine studies involving LLMs in healthcare that have conducted human evaluations. We analyze the metrics used, assess their variability across studies. We also propose a standardized framework along with an interactive open web application HumanELY, to facilitate human evaluation. We believe that use of HumanELY will provide an opportunity for consistent, comprehensive, reliable, reproducible, and measurable human evaluations of LLM in healthcare. HumanELY is publicly available at https://www.brainxai.com/humanely.

Cite

CITATION STYLE

APA

Awasthi, R., Bhattad, A., Ramachandran, S. P., Mishra, S., Khanna, A. K., Cywinski, J. B., … Mathur, P. (2025). Human evaluation of large language models in healthcare: gaps, challenges, and the need for standardization. Npj Health Systems, 2(1). https://doi.org/10.1038/s44401-025-00043-2

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free