Evaluation of the readability, quality, and accuracy of AI chatbot responses to questions about deleterious oral habits

3Citations
Citations of this article
40Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Background: Artificial intelligence (AI)-based chatbots are increasingly used by parents as convenient and fast-access sources of information on health-related topics. This study aimed to assess the readability, accuracy and overall quality of responses provided by ChatGPT-4o, Google Gemini and Microsoft Copilot to questions concerning deleterious oral habits in children. Methods: A total of 43 questions, derived from real-life discussions on the Reddit platform, were revised for clarity and demographic diversity. These were classified into seven categories based on specific types of deleterious oral habits, including thumb sucking, bruxism, pacifier use, bruxism, tongue thrusting, lip sucking, nail biting, and mouth breathing. Responses from each AI chatbot were evaluated using multiple evaluation tools including Flesch Reading Ease (FRE), Flesch-Kincaid Grade Level (FKGL), the modified DISCERN tool (mDISCERN), Global Quality Score (GQS), and misinformation scoring system. Statistical analyses were performed using the Kruskal–Wallis test followed by Dunn’s post hoc test for non-normally distributed variables, and one-way ANOVA with Tukey’s post hoc test for normally distributed variables (p

Cite

CITATION STYLE

APA

Özdemir, Ö. T., Kavan, M. Y., & Güven, Y. (2025). Evaluation of the readability, quality, and accuracy of AI chatbot responses to questions about deleterious oral habits. BMC Oral Health, 25(1). https://doi.org/10.1186/s12903-025-07298-z

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free