Large language model-generated clinical summaries in emergency departments: A blinded comparison study

0Citations
Citations of this article
2Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Emergency department (ED) clinicians routinely construct “one-liner” summaries, distilling a patient's history and presentation into one high-yield sentence that supports rapid decision-making. Producing these summaries is cognitively demanding. Large language models (LLMs) may assist by synthesizing longitudinal electronic health record (EHR) data. In this blinded study of 99 ED encounters, emergency physicians evaluated paired LLM- and physician-authored summaries on accuracy, completeness, and clinical utility, indicating their overall preference with free-text explanation. LLM-generated one-liner summaries were produced using k-nearest-neighbor few-shot prompting. We used linear mixed-effects to compare ratings. We also examined the LLM's selective note inclusion and used rapid content analysis to summarize free-text explanations. We found that across all dimensions, LLM-generated summaries received higher ratings than physician-authored summaries. Mean (SE) estimated marginal means for accuracy were 4.18 (0.09) vs 3.40 (0.11) (β = 0.78; 95% CI 0.50–1.07), for completeness 3.69 (0.10) vs 3.25 (0.12) (β = 0.44; 95% CI 0.14–0.74), and for clinical utility 3.88 (0.10) vs 3.21 (0.12) (β = 0.67; 95% CI 0.35–0.99). The LLM utilized 53.2% of available notes, with History & Physical (H&P) notes requested in 97.0% of cases and imaging reports in 81.8%, demonstrating context-sensitive document selection that varied significantly by chief complaint. Qualitative analysis indicated that LLM summaries were often more inclusive and neutrally phrased, whereas physician summaries exhibited greater contextual nuance but occasionally omitted key details. These findings represent an important first step in evaluating whether LLMs can produce clinically acceptable one-liner summaries from longitudinal EHR data; prospective validation is required before deployment in real-world ED workflows.

Cite

CITATION STYLE

APA

Golchini, N., Mehandru, N., Alaa, A., & Molina, M. (2026). Large language model-generated clinical summaries in emergency departments: A blinded comparison study. PLOS Digital Health, 5(7). https://doi.org/10.1371/journal.pdig.0001491

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free