AI at the bedside: Randomised controlled trial of ChatGPT’s impact on student performance in real-patient clinical exams

1Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Background: Generative artificial intelligence (AI) tools are entering clinical training faster than curricula and assessments can adapt. It is currently unknown whether point-of-care access to large language models (LLMs) improves clinician performance during real-time bedside assessments. Objective: To evaluate the effect of allowing ChatGPT use on student performance in ward-based real-patient clinical exams. Methods: We conducted a parallel‑group, randomised controlled trial (2:1 allocation) across four academic hospitals in a middle-income country setting. Final‑year medical students completed a 30‑minute uninterrupted patient encounter followed by a 20‑minute assessor-led evaluation. Intervention: ChatGPT (GPT‑4o) permitted during the encounter. With only minimal training, participants’ point-of-care ChatGPT use reflected self-designed approaches. Control: no digital aids. Primary outcome: overall clinical performance (0–100) scored on a standard rubric (history, examination, differential, diagnosis, investigations, management, counselling). Secondary outcomes: domain sub‑scores; observer‑rated patient interaction; student experience; patient satisfaction; subsequent summative exam scores. Analyses used ANCOVA adjusting for prior academic performance and site (α = 0.05). Results: Seventy‑three students were analysed (ChatGPT n = 49; control n = 24). Overall performance score did not differ (66.7 ± 9.9 vs 68.2 ± 7.8; unadjusted p = 0.50; adjusted p = 0.21). Domain scores showed small or negligible effects throughout. ChatGPT use patterns (frequency, duration, perceived helpfulness) were not associated with performance. Participants in the intervention arm found ChatGPT helpful overall (85%), particularly for differential diagnosis (92%) and management planning (81%), but performance gains were inconsistent; 37% reported distraction. Patients expressed high acceptance and satisfaction with student ChatGPT use. Prior academic performance significantly predicted assessment scores (p = 0.04), with no preferential benefit for weaker students. Group performance in a post-study summative clinical test was similar. Conclusions: Minimally trained, self‑directed point‑of‑care ChatGPT use did not improve bedside performance. Any benefit is likely to depend on structured training, consistent prompts or scaffolds, and clearer workflow integration. LLM integration can support, not substitute, foundational clinical competence.

Cite

CITATION STYLE

APA

Saloojee, H., Gramanie, M. C., Mwali, R., Bassett, B. A., Madhi, S. A., & Kala, I. S. (2026). AI at the bedside: Randomised controlled trial of ChatGPT’s impact on student performance in real-patient clinical exams. Medical Teacher. https://doi.org/10.1080/0142159X.2026.2652061

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free