Stroke Prediction Using the Trust Evaluation with Data Leakage Avoiding

4Citations
Citations of this article
15Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Early detection of the severe disease - stroke is a key step toward effective treatment. Stroke disease data is imbalanced and normally contains the majority of negative cases (without stroke) and the minority of positive cases (stroke). Previous work has used SMOTE to deal with imbalanced data, but most researchers have implemented it for the entire dataset, which means the "answer"was silently "be told"and saved in the entire data, causing data leakage. Moreover, the previous work uses accuracy only as the metrics make the result less guaranteed. We propose a method using the SMOTE applied to the training set only and apply 13 machine learning classifiers for predicting stroke. We combine the AUC with accuracy as evaluation metrics in the stroke prediction task, which can elevate the confidence level of the assessment of results. The experiment shows that misused SMOTE and standardization can cause data leakage and the combined metrics can evaluate models with higher trustworthiness. We conclude that using our method can avoid data leakage and assess the model with higher trustworthiness.

Cite

CITATION STYLE

APA

Ye, X., Xu, W., Ye, X., Long, D., Yin, Q., & Huang, B. (2023). Stroke Prediction Using the Trust Evaluation with Data Leakage Avoiding. In Journal of Physics: Conference Series (Vol. 2560). Institute of Physics. https://doi.org/10.1088/1742-6596/2560/1/012051

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free