Assessing the Utility of Large Language Models in Guiding Dental Practitioners on Pediatric Patient Care: A Comparative AI Study

2Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.

Abstract

Background: Large Language Models (LLMs) are transforming clinical decision-making by offering rapid, con text-aware access to evidence-based knowledge. However, their efficacy in pediatric dentistry remains underexplo red, especially across multiple LLM platforms. Objective: To comparatively evaluate the clinical quality, readability, and originality of responses generated by nine contemporary LLMs for pediatric dental queries. Material and Methods: A cross-sectional study assessed the performance of ChatGPT-3.5, ChatGPT-4o, Gemini 2.0, Gemini 2.5, Claude 3.5 Haiku, Claude 3.7 Sonnet, Grok-3, Grok-3 Mini, and DeepSeek-V3. Twenty pediatric dental questions were posed in one-shot queries to each LLM. Responses were evaluated by ten pediatric dental experts using the Modified Global Quality Scale (MGQS), Flesch Reading Ease Score (FRES), Flesch-Kincaid Grade Level (FKGL), and Turnitin Similarity Index. ANOVA and Cohen’s Kappa were used for statistical analysis. Results: ChatGPT-4o demonstrated the highest overall MGQS (4.28 ± 0.24), followed by ChatGPT-3.5 (3.45 ± 0.27). DeepSeek-V3 scored lowest (2.18 ± 0.19). Topic-wise, ChatGPT-4o consistently outperformed others across all subdomains. FRES and FKGL scores indicated moderate readability, with Claude models exhibiting the highest linguistic complexity. Turnitin analysis revealed low-to-moderate similarity across models. Inter-rater agreement was substantial (κ = 0.78). Conclusions: Among evaluated LLMs, ChatGPT-4o exhibited superior performance in clinical relevance, coherence, and originality, suggesting its potential utility as an adjunct in pediatric dental decision-making. Nonetheless, variabi lity across models underscores the need for critical appraisal and cautious integration into clinical workflows.

Cite

CITATION STYLE

APA

Raj, M., Ravindran, V., & Arthanari, A. (2025). Assessing the Utility of Large Language Models in Guiding Dental Practitioners on Pediatric Patient Care: A Comparative AI Study. Journal of Clinical and Experimental Dentistry, 17(9). https://doi.org/10.4317/jced.63136

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free