Performance of ChatGPT on the Chinese Postgraduate Examination for Clinical Medicine: Survey Study

Peng Yu; Changchang Fang; Xiaolin Liu; Wanying Fu; Jitao Ling; Zhiwei Yan; Yuan Jiang; Zhengyu Cao; Maoxiong Wu; Zhiteng Chen; Wengen Zhu; Yuling Zhang; Ayiguli Abudukeremu; Yue Wang; Xiao Liu; Jingfeng Wang

Journal ArticleOPEN ACCESS

Performance of ChatGPT on the Chinese Postgraduate Examination for Clinical Medicine: Survey Study

JMIR Medical Education (2024) 10(1)

DOI: 10.2196/48514

5Citations

54Readers

Get full text

Abstract

Background: ChatGPT, an artificial intelligence (AI) based on large-scale language models, has sparked interest in the field of health care. Nonetheless, the capabilities of AI in text comprehension and generation are constrained by the quality and volume of available training data for a specific language, and the performance of AI across different languages requires further investigation. While AI harbors substantial potential in medicine, it is imperative to tackle challenges such as the formulation of clinical care standards; facilitating cultural transitions in medical education and practice; and managing ethical issues including data privacy, consent, and bias. Objective: The study aimed to evaluate ChatGPT's performance in processing Chinese Postgraduate Examination for Clinical Medicine questions, assess its clinical reasoning ability, investigate potential limitations with the Chinese language, and explore its potential as a valuable tool for medical professionals in the Chinese context. Methods: A data set of Chinese Postgraduate Examination for Clinical Medicine questions was used to assess the effectiveness of ChatGPT's (version 3.5) medical knowledge in the Chinese language, which has a data set of 165 medical questions that were divided into three categories: (1) common questions (n=90) assessing basic medical knowledge, (2) case analysis questions (n=45) focusing on clinical decision-making through patient case evaluations, and (3) multichoice questions (n=30) requiring the selection of multiple correct answers. First of all, we assessed whether ChatGPT could meet the stringent cutoff score defined by the government agency, which requires a performance within the top 20% of candidates. Additionally, in our evaluation of ChatGPT's performance on both original and encoded medical questions, 3 primary indicators were used: accuracy, concordance (which validates the answer), and the frequency of insights. Results: Our evaluation revealed that ChatGPT scored 153.5 out of 300 for original questions in Chinese, which signifies the minimum score set to ensure that at least 20% more candidates pass than the enrollment quota. However, ChatGPT had low accuracy in answering open-ended medical questions, with only 31.5% total accuracy. The accuracy for common questions, multichoice questions, and case analysis questions was 42%, 37%, and 17%, respectively. ChatGPT achieved a 90% concordance across all questions. Among correct responses, the concordance was 100%, significantly exceeding that of incorrect responses (n=57, 50%; P

Author supplied keywords

Cite

CITATION STYLE

APA

Yu, P., Fang, C., Liu, X., Fu, W., Ling, J., Yan, Z., … Wang, J. (2024). Performance of ChatGPT on the Chinese Postgraduate Examination for Clinical Medicine: Survey Study. JMIR Medical Education, 10(1). https://doi.org/10.2196/48514

Performance of ChatGPT on the Chinese Postgraduate Examination for Clinical Medicine: Survey Study

Abstract

Author supplied keywords

Cite

Register to see more suggestions