Abstract
Obesity has become a major global health challenge, with prevalence rising across all age groups. While Body Mass Index (BMI) is widely used for diagnosis, it has limitations, as it cannot fully capture lifestyle or behavioral risk factors. This study applies machine learning (ML) techniques to classify obesity categories using the UCI dataset “Estimation of Obesity Levels Based on Eating Habits and Physical Condition.” The dataset includes 16 lifestyle and demographic features, and experiments were conducted with and without height and weight to evaluate the potential effect of BMI-related variables. Random Forest was adopted as the primary model, with additional classifiers employed as comparative methods. Results show that ensemble tree-based models achieved the highest performance overall. In particular, Random Forest demonstrated strong accuracy even in the absence of height and weight, suggesting that it can capture meaningful lifestyle and behavioral patterns without relying solely on BMI. Feature importance analysis identified age, diet, and physical activity as key predictors. Misclassification analysis further revealed that some severely obese individuals were incorrectly predicted as normal weight, especially those with positive family history but healthier reported habits. These findings highlight the potential of ML to support obesity risk assessment and the importance of addressing information leakage when BMI-related features are included.
Cite
CITATION STYLE
Zhang, J. (2025). Machine Learning Analysis of Obesity Levels Based on Lifestyle and Physical Features. Theoretical and Natural Science, 135(1), 90–98. https://doi.org/10.54254/2753-8818/2025.au27039
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.