Abstract
Customer churn is a critical challenge for subscription-based businesses, often exacerbated by imbalanced datasets that hinder predictive accuracy. This study evaluates various oversampling techniques, K-means SMOTE, SMOTE, and ADASYN, that generate synthetic samples to balance datasets. The objective is to assess the impact of these oversampling techniques on the performance of machine learning (ML) classifiers, including gradient boosting (GB), random forest (RF), naive Bayes (NB), and support vector machines (SVM). Findings reveal that K-means SMOTE is the most effective at improving model performance, while GB consistently outperforms other classifiers in churn prediction. These findings provide valuable insights into optimizing data balancing and predictive models, offering a robust framework to enhance customer retention strategies.
Author supplied keywords
Cite
CITATION STYLE
Abdalla, F. A., Mohammed, Z. M. S., Satty, A., Mahmoud, A. F. A., Ammar, M. B., Mohamed, A. S., … Ahmed, S. A. (2025). Comparative Analysis of Data Augmentation Methods for Enhancing the Performance of Churn Prediction Models. International Journal of Analysis and Applications, 23. https://doi.org/10.28924/2291-8639-23-2025-296
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.