Performance of Machine Learning Algorithms for Credit Risk Prediction with Feature Selection

1Citations
Citations of this article
11Readers
Mendeley users who have this article in their library.

Abstract

Financial institutions increasingly rely on machine learning (ML) models to assess credit risk and make lending decisions. Accurate prediction hinges on effective feature selection, which can significantly enhance model performance. This paper investigates the efficacy of seven supervised ML algorithms in predicting credit risk: Naive Bayes, Support Vector Machine, Decision Tree, K-Nearest Neighbor, Artificial Neural Network, Random Forest, and Logistic Regression. Using a German credit dataset comprising 1000 observations with 20 explanatory variables, we evaluated model performance using accuracy, kappa statistic, and F1 score. Two data-splitting scenarios (70-30% and 80-20%) were employed to assess robustness. We addressed outliers through imputation methods to optimize model performance and applied the Boruta algorithm for feature selection, which identified and eliminated six non-contributing features. Our findings consistently demonstrate the superiority of the Random Forest algorithm across both scenarios. Regarding accuracy, Random Forest achieved 77.3% in the 70-30% split and 80% in the 80-20% split, outperforming all other methods. These results underscore the potential of Random Forest as a valuable tool for credit risk assessment in financial institutions.

Cite

CITATION STYLE

APA

Seliem, M. M., Amin, M., El Nasr, M. M. A., Elnaggar, E. A., Khalifa, H. A. M., & Arab, M. A. A. (2025). Performance of Machine Learning Algorithms for Credit Risk Prediction with Feature Selection. Statistics, Optimization and Information Computing, 14(1), 311–328. https://doi.org/10.19139/soic-2310-5070-2392

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free