Predicting and Explaining Middle-School Dropout Risk on Imbalanced Data

0Citations
Citations of this article
2Readers
Mendeley users who have this article in their library.

Abstract

Early identification of students at risk of dropping out is vital for providing timely support and efficiently allocating educational resources. Using a large, real-world middle-school cohort (N = 810,853; prevalence 4.38%), we develop an explainable TabNet-based model for tabular data. We train with class-weighted loss and early stop on validation PR-AUC to prioritize minority detection, we calibrate probabilities via Platt scaling, and fix a single operating threshold by maximizing Youdena s J. The model achieves strong discrimination (ROC-AUC ≈ 0.90; PR-AUC ≈ 0.47 vs. 0.043 baseline) and a recall-centric operating point on test (TPR = 0.793, TNR = 0.833), with balanced metrics confirming robustness (G-Mean ≈ 0.813; MCC ≈0.324). Calibration markedly improves probability quality (Brier score 0.1234 ≈ 0.0317), to better reflect the true likelihood of outcomes. The use of the mask-based TabNet feature provides transparent reasons where the most recent cumulative grade point average, academic delay, and socioeconomic vulnerability/poverty dominate, supporting targeted and interpretable intervention. The study also reveals that some students are flagged as dropouts despite continuing, necessitating real follow-up costs.

Cite

CITATION STYLE

APA

El Jihaoui, M., Abra, O. E. K., & Mansouri, K. (2025). Predicting and Explaining Middle-School Dropout Risk on Imbalanced Data. In E3S Web of Conferences (Vol. 680). EDP Sciences. https://doi.org/10.1051/e3sconf/202568000073

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free