Abstract
Background: The deployment of OpenAI's ChatGPT 3.5 and its subsequent versions, ChatGPT 4 and 4 with Vision (4V), has notably influenced the medical field. Demonstrating remarkable performance in medical exams globally, these models show potential for educational applications. However, their effectiveness in non-English contexts, particularly in Chile's Medical Licensing Exam, a critical step for medical practitioners in Chile, is less explored. This gap highlights the need to evaluate ChatGPT's adaptability to diverse linguistic and cultural contexts. Objective: This study aims to evaluate the proficiency of ChatGPT versions 3.5, 4, and 4V in answering EUNACOM (Examen Único Nacional de Conocimientos de Medicina), a major medical examination in Chile. Methods: Three official practice drills (540 questions) from the University of Chile, mirroring the EUNACOM structure and difficulty, were used to test ChatGPT versions 3.5, 4, and 4V. The three ChatGPT versions underwent three attempts of answering each drill. Responses to questions during each round were systematically categorized and analyzed to assess the accuracy rate of the responses. Results: All versions of ChatGPT passed the EUNACOM drills, with version 4 and version 4V outperforming version 3.5 (P
Author supplied keywords
Cite
CITATION STYLE
Pino, M. R., Duarte, M. R., Jaramillo, V. B., Pérez, J. T., & Salehi, S. (2024). Exploring the Proficiency of ChatGPT 3.5, 4, and 4 with Vision in the Chilean Medical Licensing Exam: An Observational Study. JMIR Medical Education, 10. https://doi.org/10.2196/55048
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.