Abstract
Early identification of students at risk of dropping out is vital for providing timely support and efficiently allocating educational resources. Using a large, real-world middle-school cohort (N = 810,853; prevalence 4.38%), we develop an explainable TabNet-based model for tabular data. We train with class-weighted loss and early stop on validation PR-AUC to prioritize minority detection, we calibrate probabilities via Platt scaling, and fix a single operating threshold by maximizing Youdena s J. The model achieves strong discrimination (ROC-AUC ≈ 0.90; PR-AUC ≈ 0.47 vs. 0.043 baseline) and a recall-centric operating point on test (TPR = 0.793, TNR = 0.833), with balanced metrics confirming robustness (G-Mean ≈ 0.813; MCC ≈0.324). Calibration markedly improves probability quality (Brier score 0.1234 ≈ 0.0317), to better reflect the true likelihood of outcomes. The use of the mask-based TabNet feature provides transparent reasons where the most recent cumulative grade point average, academic delay, and socioeconomic vulnerability/poverty dominate, supporting targeted and interpretable intervention. The study also reveals that some students are flagged as dropouts despite continuing, necessitating real follow-up costs.
Cite
CITATION STYLE
El Jihaoui, M., Abra, O. E. K., & Mansouri, K. (2025). Predicting and Explaining Middle-School Dropout Risk on Imbalanced Data. In E3S Web of Conferences (Vol. 680). EDP Sciences. https://doi.org/10.1051/e3sconf/202568000073
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.