Abstract
Financial institutions increasingly rely on machine learning (ML) models to assess credit risk and make lending decisions. Accurate prediction hinges on effective feature selection, which can significantly enhance model performance. This paper investigates the efficacy of seven supervised ML algorithms in predicting credit risk: Naive Bayes, Support Vector Machine, Decision Tree, K-Nearest Neighbor, Artificial Neural Network, Random Forest, and Logistic Regression. Using a German credit dataset comprising 1000 observations with 20 explanatory variables, we evaluated model performance using accuracy, kappa statistic, and F1 score. Two data-splitting scenarios (70-30% and 80-20%) were employed to assess robustness. We addressed outliers through imputation methods to optimize model performance and applied the Boruta algorithm for feature selection, which identified and eliminated six non-contributing features. Our findings consistently demonstrate the superiority of the Random Forest algorithm across both scenarios. Regarding accuracy, Random Forest achieved 77.3% in the 70-30% split and 80% in the 80-20% split, outperforming all other methods. These results underscore the potential of Random Forest as a valuable tool for credit risk assessment in financial institutions.
Author supplied keywords
Cite
CITATION STYLE
Seliem, M. M., Amin, M., El Nasr, M. M. A., Elnaggar, E. A., Khalifa, H. A. M., & Arab, M. A. A. (2025). Performance of Machine Learning Algorithms for Credit Risk Prediction with Feature Selection. Statistics, Optimization and Information Computing, 14(1), 311–328. https://doi.org/10.19139/soic-2310-5070-2392
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.