Abstract
Introduction: To evaluate the accuracy and completeness of responses across common obstetrical and gynecologic topics generated by the large language models (LLMs) ChatGPT and Google Gemini, which have become increasingly popular for patients seeking medical information before physician consultations. Methods: Ten topics were identified, five obstetrical (prenatal labs, extended carrier screen, treatments for nausea and vomiting in pregnancy, gestational diabetes, and trial of labor after cesarean section) and five gynecologic (polycystic ovary syndrome, pelvic inflammatory disease, cervical smears, mammograms, and birth control). For each condition, ChatGPT generated five of the most frequently asked patient questions, which were then presented separately to ChatGPT and Google Gemini. Board-certified Obstetrics and Gynecology physicians evaluated the responses using Likert scales for accuracy (1–6) and completeness (1–3). Results: Acceptable response criteria were defined as an accuracy score of 5 or greater (“nearly all correct”) and a completeness score of 2 or greater (“adequately complete”). Most responses from both models met these thresholds. Wilcoxon signed-rank tests demonstrated statistically significant differences in accuracy and completeness between models (P < 0.05). Inter-rater agreement was measured using intraclass correlation coefficients. For obstetrical topics, ChatGPT scored −0.047 (completeness) and 0.112 (accuracy), whereas Google Gemini scored 0.367 and 0.205, respectively. For gynecologic topics, ChatGPT scored 0.328 and 0.20, compared with Google Gemini at 0.151 and −0.08. Conclusion: Both LLMs provided largely accurate and complete responses to patient questions. ChatGPT demonstrated stronger outcomes overall, suggesting potential utility in patient education; however, patients should confirm online information with physicians given the limitations of LLMs.
Author supplied keywords
Cite
CITATION STYLE
West, M., Alsaidi, A., Siddiqi, R., Sayyed, F., Counts, R., Quinto, L., & Stansbury, N. (2026). Artificial intelligence in obstetrics and gynecology: Evaluating ChatGPT and Google Gemini in answering patient questions. International Journal of Gynecology and Obstetrics, 173(1), 328–335. https://doi.org/10.1002/ijgo.70622
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.