Comparative evaluation of ChatGPT, Gemini, and Grok in clinical decision-making and general knowledge assessment for impacted maxillary canines

0Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

Objective: This study aimed to compare extraction versus orthodontic eruption decisions for impacted maxillary canines made by three artificial intelligence-based chatbots (ChatGPT, Gemini, and Grok) with those made by orthodontist raters, and to evaluate the overall accuracy of these artificial intelligence-generated recommendations. Methods: Thirty-three patients with impacted maxillary canines were selected, and standardized case scenarios incorporating key diagnostic parameters were presented to the three chatbots. Their treatment decisions were recorded and compared with orthodontists’ consensus decisions. Additionally, 10 general queries regarding impacted maxillary canines were submitted to the chatbots. The responses were rated by three orthodontists using a modified 5-point Global Quality Score. Results: The chatbots and orthodontists showed moderate agreement regarding treatment decisions (κ=0.411–0.524, P < 0.05). Gemini produced significantly more discordant responses, frequently over-recommending orthodontic eruptions (P=0.002), whereas Grok and ChatGPT received significantly higher scores than Gemini in the case-based scenarios (P < 0.001). Grok outperformed both ChatGPT and Gemini for general queries (P=0.006). Conclusions: While Gemini showed lower clinical alignment with orthodontists for treatment decisions regarding impacted canines, ChatGPT and Grok demonstrated moderate agreement with orthodontists and produced relatively accurate responses. These findings highlight the potential of chatbots as supportive tools for orthodontic decision-making. However, their use requires careful supervision to avoid the risks associated with inaccurate or misleading recommendations.

Cite

CITATION STYLE

APA

Sabah, G. A., & Kanmaz, M. G. (2026). Comparative evaluation of ChatGPT, Gemini, and Grok in clinical decision-making and general knowledge assessment for impacted maxillary canines. Korean Journal of Orthodontics, 56(1), 45–56. https://doi.org/10.4041/kjod25.174

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free