Assessing the proficiency of large language models on funduscopic disease knowledge

0Citations
Citations of this article
2Readers
Mendeley users who have this article in their library.

Abstract

AIM: To assess the performance of five distinct large language models (LLMs; ChatGPT-3.5, ChatGPT-4, PaLM2, Claude 2, and SenseNova) in comparison to two human cohorts (a group of funduscopic disease experts and a group of ophthalmologists) on the specialized subject of funduscopic disease. METHODS: Five distinct LLMs and two distinct human groups independently completed a 100-item funduscopic disease test. The performance of these entities was assessed by comparing their average scores, response stability, and answer confidence, thereby establishing a basis for evaluation. RESULTS: Among all the LLMs, ChatGPT-4 and PaLM2 exhibited the most substantial average correlation. Additionally, ChatGPT-4 achieved the highest average score and demonstrated the utmost confidence during the exam. In comparison to human cohorts, ChatGPT-4 exhibited comparable performance to ophthalmologists, albeit falling short of the expertise demonstrated by funduscopic disease specialists. CONCLUSION: The study provides evidence of the exceptional performance of ChatGPT-4 in the domain of funduscopic disease. With continued enhancements, validated LLMs have the potential to yield unforeseen advantages in enhancing healthcare for both patients and physicians.

Cite

CITATION STYLE

APA

Wu, J. Y., Zeng, Y. M., Qian, X. Z., Hong, Q., Hu, J. Y., Wei, H., … Shao, Y. (2025). Assessing the proficiency of large language models on funduscopic disease knowledge. International Journal of Ophthalmology, 18(7), 1205–1213. https://doi.org/10.18240/ijo.2025.07.03

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free