Reliability and Readability Assessment of Atrial Fibrillation Patient Information Delivered by Artificial Intelligence-Based Language Models (ChatGPT, YouChat, Gemini, and Perplexity AI) in English and Spanish

Emilio Jose Juan Juan-Guardela; Jesús Andrés Beltrán-España; María Paula Ravagli-Baquero; Cristian Orlando Porras-Bueno; Edward Cáceres-Méndez; Daniel Fernandez Ávila; Oscar Muñoz-Velandia; Ángel Alberto García-Peña

Journal ArticleOPEN ACCESS

Reliability and Readability Assessment of Atrial Fibrillation Patient Information Delivered by Artificial Intelligence-Based Language Models (ChatGPT, YouChat, Gemini, and Perplexity AI) in English and Spanish

Clinical Medicine Insights: Cardiology (2025) 19

DOI: 10.1177/11795468251383666

0Citations

5Readers

Abstract

Background: Atrial fibrillation (AF) is the most prevalent arrhythmia and a significant cause of morbidity. Artificial intelligence (AI)-based language models represent a novel tool for searching for medical information; however, there is still uncertainty regarding their reliability and readability in different languages. Objective: To assess the reliability and readability of information provided by AI-based models for patients with AF. Methods: A cross-sectional study was conducted to assess the reliability and readability of the responses generated by ChatGPT, YouChat, Gemini and Perplexity on AF in English and Spanish. Thirty standardised questions were posed in both languages. The quality of the responses was then assessed by 2 independent reviewers via a standardised tool. Readability was assessed via the Flesch–Szigrist formula. The results were then compared by tool and language. Results: ChatGPT demonstrated the highest interrater agreement (PA = 0.73 in Spanish, 0.80 in English), followed by Gemini in English (PA = 0.66). In Spanish, ChatGPT generated the highest percentage of complete responses (80%), followed by Perplexity (73%) and Gemini (47%). In English, Perplexity demonstrated the strongest performance, with a score of 93%, followed by ChatGPT, with 73%, and Gemini, with 53%. A readability analysis revealed significant differences between the models (P < .01). The ChatGPT demonstrated the highest performance, although its content was moderately challenging in Spanish and highly challenging in English. Conclusion: ChatGPT and Perplexity emerged as the most reliable models, although readability remains a concern. There is a clear need for improvements to optimise the accuracy and accessibility of AI-generated medical information.

Author supplied keywords

Cite

CITATION STYLE

APA

Juan-Guardela, E. J. J., Beltrán-España, J. A., Ravagli-Baquero, M. P., Porras-Bueno, C. O., Cáceres-Méndez, E., Ávila, D. F., … García-Peña, Á. A. (2025). Reliability and Readability Assessment of Atrial Fibrillation Patient Information Delivered by Artificial Intelligence-Based Language Models (ChatGPT, YouChat, Gemini, and Perplexity AI) in English and Spanish. Clinical Medicine Insights: Cardiology, 19. https://doi.org/10.1177/11795468251383666

Reliability and Readability Assessment of Atrial Fibrillation Patient Information Delivered by Artificial Intelligence-Based Language Models (ChatGPT, YouChat, Gemini, and Perplexity AI) in English and Spanish

Abstract

Author supplied keywords

Cite

Register to see more suggestions