Abstract
Objective: Inpatients undergoing stroke rehabilitation experience high malnutrition rates, requiring strict dietary management. However, manual and time-pressured dietary provision can cause errors in diet composition, highlighting the need for innovation. Therefore, we aimed to evaluate whether GPT-4o can accurately identify dietary errors in hospital-based stroke rehabilitation menus, analyze differences in AI vs. expert rationale for decisions, and explore AI’s potential role in clinical workflows through a structured collaboration framework. Methods: A TRIPOD-compliant validation study analyzing 264 hospital-based menus designed for stroke rehabilitation inpatients requiring specialized diets (e.g., dysphagia, diabetes). GPT-4o’s dietary compliance classifications were assessed using a structured 0-error, 1-error, and 2+ error framework, with expert dietitians as ground-truth in a rehabilitation hospital nutrition department, where expert dietitians selected menus from existing clinical practices for inpatients on specialized diets. AI-expert agreement, overall accuracy, sensitivity, and specificity in dietary error classification were assessed. AI vs. expert justifications were analyzed thematically to identify differences in decision rationale. Cohen’s Kappa (95% CI) measured inter-rater reliability. Overall accuracy, sensitivity, and specificity were calculated using a 3 × 3 confusion matrix, comparing AI classifications (0-error, 1-error, 2+ error) to the expert-labeled ground truth. Thematic analysis categorized AI vs. expert justifications for flagged dietary errors. Results: Out of 264 menus (1,000+ food items), 26 (9.8%) had discrepancies. Among these, 57.7% (15 cases) were PAS-based dysphagia diets, followed by diabetic (19.2%, 5 cases) and allergen-related (15.4%, 4 cases) diets. The remaining two cases involved low-sodium and low-fat diets. Cohen’s Kappa: 0.892 (95% CI: 0.845–0.939, p < 0.001). 0-errors: Sensitivity 94.3%, specificity 100%; 1-error: Sensitivity 86.2%, specificity 96.6%; 2+-errors: Sensitivity 97.8%, specificity 92.6%. Thematic analysis revealed GPT-4o followed strict rule-based interpretations, whereas dietitians incorporated patient tolerance and food preparation considerations. Conclusion: GPT-4o demonstrated high accuracy but over-flagged violations, supporting its role as a prescreening tool with expert collaboration.
Author supplied keywords
Cite
CITATION STYLE
García-Rudolph, A., Hernandez-Pena, E., del Cacho, N., Teixidó-Font, C., Wright, M. A., & Opisso, E. (2026). GPT-4o in Nutrition for Inpatients Undergoing Post-Stroke Rehabilitation: Identifying Dietary Errors, Exploring Expert-AI Rationale Differences, and Structuring AI-Expert Collaboration. Journal of the American Nutrition Association, 45(4), 310–321. https://doi.org/10.1080/27697061.2025.2571878
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.