A Comparative Analysis of Six Machine Learning Classifiers for Early-Stage Diabetes Risk Prediction Using a Kaggle Dataset

1Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Diabetes is a common metabolic condition characterized by an elevated blood sugar level due to impaired insulin production or action. Adverse sequelae of diabetes may be kidney damage, neuropathy, cardiovascular disease, and eye problems. Diabetes is increasingly becoming a regular phenomenon across the globe, and so, averting its impact on individuals and the healthcare systems will be to carry out early diagnosis of the disease, proper curative therapy, and preventive strategies. A study comparing various machine learning (ML) classifiers, including K-Nearest Neighbors (KNN), random forests (RF), Logistic Regression (LR), Gradient Boosting (GB), XGBoost, and decision trees (DT), was conducted to estimate the likelihood of diabetes. The model is evaluated by calculating accuracy, precision, recall, F1-score, execution time, and confusion matrix analysis. With the highest F1-score (0.99), accuracy (0.99), and recall (0.99), the Random Forest classifier performed exceptionally well, exhibiting remarkable resilience and classification capability. The accuracy, recall, and F1-score of both GB and XGBoost were 0.97, 0.96, and 0.97, respectively; however, XGBoost’s execution time was longer than GB’s. The decision tree model outperformed the LR model, achieving an accuracy of 0.92, a recall of 0.96, and an F1-score of 0.94. The decision tree model had an accuracy of 0.95, a recall of 0.93, and an F1-score of 0.95. The KNN model’s accuracy, recall, and F1-score were 0.90, 0.89, and 0.93, respectively. With both high prediction accuracy and high sensitivity to positive cases, Random Forest is the best model for predicting diabetes overall, according to the data. This study makes it a good choice for applications needing early detection.

Cite

CITATION STYLE

APA

Asaad, A. M., Qahtan, M. H., & Younis, A. K. (2025). A Comparative Analysis of Six Machine Learning Classifiers for Early-Stage Diabetes Risk Prediction Using a Kaggle Dataset. Ingenierie Des Systemes d’Information, 30(12), 3235–3241. https://doi.org/10.18280/isi.301216

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free