A comparative study on the application of large language models: Deepseek-R1, GPT-4o, and Claude-Sonnet-4 in post-cardiac surgery rehabilitation—A cross-sectional study

2Citations
Citations of this article
23Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Objective: This study aimed to systematically evaluate the performance of three advanced Chinese large language models (LLMs)—DeepSeek-R1, GPT-4o, and Claude-Sonnet-4—in supporting patient education during post-cardiac surgery rehabilitation. Methods: A total of 35 patient-centered questions were developed based on clinical guidelines, covering five core domains: postoperative care, medication and diet, mental health, complication prevention, and physical activity. Each model was prompted with the same questions five times. Responses were independently assessed by clinical experts for accuracy, completeness, readability (using FRE and FKGL), and reproducibility under repeated prompting. Statistical analyses were conducted using analysis of variance and post-hoc least significant difference (LSD) tests. Results: DeepSeek-R1 demonstrated the highest overall performance in terms of accuracy (mean score: 4.64) and completeness (4.20), and achieved the highest response stability (85.7%). GPT-4o outperformed the others in readability (FRE: 53.19) and linguistic fluency but showed lower reproducibility (62.9%). Claude-Sonnet-4 showed moderate and variable performance, with limitations in clinical detail. All observed differences were statistically significant (P < 0.05). Conclusion: DeepSeek-R1 is most suitable for structured, guideline-based rehabilitation education. GPT-4o may be preferable in emotionally supportive or patient-facing scenarios due to its superior readability. Claude-Sonnet-4, while less consistent, may offer stylistic balance in diverse communication settings. Overall, LLMs show promising potential in digital cardiac rehabilitation, but task-specific model alignment remains essential.

Cite

CITATION STYLE

APA

Li, W., & Li, Q. (2025). A comparative study on the application of large language models: Deepseek-R1, GPT-4o, and Claude-Sonnet-4 in post-cardiac surgery rehabilitation—A cross-sectional study. Digital Health, 11. https://doi.org/10.1177/20552076251393385

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free