Abstract
Despite the extensive application of machine learning (ML) methods to educational datasets, few studies have provided a systematic benchmarking of the available algorithms with respect to both predictive performance and interpretability of the resulting models. In this work, we address this gap by comparing a range of supervised learning methods on a freely available dataset concerning two high schools, where the goal is to predict student performance by modeling it as a binary classification task. Given the high feature-to-sample ratio, the problem falls within the small-data learning regime, which often challenges ML models by diluting informative features among many irrelevant ones. The experimental results show that several algorithms can achieve robust predictive performance, even in this scenario and in the presence of class imbalance. Moreover, we show how the output of ML algorithms can be interpreted and used to identify the most relevant predictors, without any a priori assumption about their impact. Finally, we perform additional experiments by removing the two most dominant features, revealing that ML models can still uncover alternative predictive patterns, thus demonstrating their adaptability and capacity for knowledge extraction under small-data conditions. Future work could benefit from richer datasets, including longitudinal data and psychological features, to better profile students and improve the identification of at-risk individuals.
Author supplied keywords
Cite
CITATION STYLE
Vecchi, E. (2026). Investigating the Efficacy and Interpretability of ML Classifiers for Student Performance Prediction in the Small-Data Regime. Education Sciences, 16(1). https://doi.org/10.3390/educsci16010149
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.