GPT-4o in Nutrition for Inpatients Undergoing Post-Stroke Rehabilitation: Identifying Dietary Errors, Exploring Expert-AI Rationale Differences, and Structuring AI-Expert Collaboration

1Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Objective: Inpatients undergoing stroke rehabilitation experience high malnutrition rates, requiring strict dietary management. However, manual and time-pressured dietary provision can cause errors in diet composition, highlighting the need for innovation. Therefore, we aimed to evaluate whether GPT-4o can accurately identify dietary errors in hospital-based stroke rehabilitation menus, analyze differences in AI vs. expert rationale for decisions, and explore AI’s potential role in clinical workflows through a structured collaboration framework. Methods: A TRIPOD-compliant validation study analyzing 264 hospital-based menus designed for stroke rehabilitation inpatients requiring specialized diets (e.g., dysphagia, diabetes). GPT-4o’s dietary compliance classifications were assessed using a structured 0-error, 1-error, and 2+ error framework, with expert dietitians as ground-truth in a rehabilitation hospital nutrition department, where expert dietitians selected menus from existing clinical practices for inpatients on specialized diets. AI-expert agreement, overall accuracy, sensitivity, and specificity in dietary error classification were assessed. AI vs. expert justifications were analyzed thematically to identify differences in decision rationale. Cohen’s Kappa (95% CI) measured inter-rater reliability. Overall accuracy, sensitivity, and specificity were calculated using a 3 × 3 confusion matrix, comparing AI classifications (0-error, 1-error, 2+ error) to the expert-labeled ground truth. Thematic analysis categorized AI vs. expert justifications for flagged dietary errors. Results: Out of 264 menus (1,000+ food items), 26 (9.8%) had discrepancies. Among these, 57.7% (15 cases) were PAS-based dysphagia diets, followed by diabetic (19.2%, 5 cases) and allergen-related (15.4%, 4 cases) diets. The remaining two cases involved low-sodium and low-fat diets. Cohen’s Kappa: 0.892 (95% CI: 0.845–0.939, p < 0.001). 0-errors: Sensitivity 94.3%, specificity 100%; 1-error: Sensitivity 86.2%, specificity 96.6%; 2+-errors: Sensitivity 97.8%, specificity 92.6%. Thematic analysis revealed GPT-4o followed strict rule-based interpretations, whereas dietitians incorporated patient tolerance and food preparation considerations. Conclusion: GPT-4o demonstrated high accuracy but over-flagged violations, supporting its role as a prescreening tool with expert collaboration.

Cite

CITATION STYLE

APA

García-Rudolph, A., Hernandez-Pena, E., del Cacho, N., Teixidó-Font, C., Wright, M. A., & Opisso, E. (2026). GPT-4o in Nutrition for Inpatients Undergoing Post-Stroke Rehabilitation: Identifying Dietary Errors, Exploring Expert-AI Rationale Differences, and Structuring AI-Expert Collaboration. Journal of the American Nutrition Association, 45(4), 310–321. https://doi.org/10.1080/27697061.2025.2571878

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free