The performance of ChatGPT-4.0o in medical imaging evaluation: a cross-sectional study

14Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.

Abstract

This study investigated the performance of ChatGPT-4.0o in evaluating the quality of positioning in radiographic images. Thirty radiographs depicting a variety of knee, elbow, ankle, hand, pelvis, and shoulder projections were produced using anthropomorphic phantoms and uploaded to ChatGPT-4.0o. The model was prompted to provide a solution to identify any positioning errors with justification and offer improvements. A panel of radiographers assessed the solutions for radiographic quality based on established positioning criteria, with a grading scale of 1-5. In only 20% of projections, ChatGPT-4.0o correctly recognized all errors with justifications and offered correct suggestions for improvement. The most commonly occurring score was 3 (9 cases, 30%), wherein the model recognized at least 1 specific error and provided a correct improvement. The mean score was 2.9. Overall, low accuracy was demonstrated, with most projections receiving only partially correct solutions. The findings reinforce the importance of robust radiography education and clinical experience.

Cite

CITATION STYLE

APA

Arruzza, E. S., Evangelista, C. M., & Chau, M. (2024). The performance of ChatGPT-4.0o in medical imaging evaluation: a cross-sectional study. Journal of Educational Evaluation for Health Professions, 21. https://doi.org/10.3352/jeehp.2024.21.29

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free