Evaluating the Effectiveness of ChatGPT Versus Human Proctors in Grading Medical Students’ Post-OSCE Notes

1Citations
Citations of this article
20Readers
Mendeley users who have this article in their library.

Abstract

Background and Objectives: Artificial intelligence (AI) tools have potential utility in multiple domains, including medical education. However, educators have yet to evaluate AI’s assessment of medical students’ clinical reasoning as evidenced in note-writing. This study compares ChatGPT with a human proctor’s grading of medical students’ notes. Methods: A total of 127 subjective, objective, assessment, and plan notes, derived from an objective structured clinical examination, were previously graded by a physician proctor across four categories: history, physical exam, differential diagnosis/thought process, and treatment plan. ChatGPT-4, using the same rubric, was tasked with evaluating these 127 notes. We compared AI-generated scores with proctors’ scores using t tests and χ2 analysis. Results: The grades assigned by ChatGPT were significantly different than those assigned by proctors in history (P

Cite

CITATION STYLE

APA

Thomas, K., Szalacha, L., Hanna, K., Anibal, J., & Petrilli, J. (2025). Evaluating the Effectiveness of ChatGPT Versus Human Proctors in Grading Medical Students’ Post-OSCE Notes. Family Medicine, 57(10), 727–731. https://doi.org/10.22454/FamMed.2025.954255

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free