Abstract
Predicting Customer churn is one of the telecommunication industry's biggest challenges. Why did their customers quit using their product, site, service, or subscription? Machine learning with Spark and Hadoop has considerably increased the ability to predict customer behaviours. The most popular predictive models, such as logistic regression, Binary Classification Evaluator, and Multi Classification Evaluator, have been used in the prediction process. Enhancing and outfit approaches are used on the training dataset to examine the impact on model effectiveness. Additionally, to further optimize the hyperparameters and produce the models, a K-fold cross-validation method is utilized to train the dataset. Finally, the test data were examined by the AUC-ROC curve and confusion matrix. In this research, an adaptation of Spark and Hadoop frameworks is made to predict customer churn. The data is pre-processed, feature analyses are performed, and the feature selection is carried out using the Vector Assembler algorithm. This study aims to analyse customer behaviors by using a dataset.
Author supplied keywords
Cite
CITATION STYLE
Verma, P., Sharma, I., Deshmukh, S., & Vashisht, R. (2023). Customer Churn Analysis using Spark and Hadoop. International Journal of Performability Engineering, 19(10), 663–675. https://doi.org/10.23940/ijpe.23.10.p4.663675
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.