Abstract
Machine learning (ML) models are frequently used to classify mental health information from textual data, but their practical use is constrained by their poor interpretability and lack of tools to fix training-related reasoning errors. Explainable AI (XAI) approaches currently in use mostly offer post hoc explanations without methodically utilizing explanation quality to enhance model performance. This paper proposes an explanation-driven iterative learning framework for classifying texts related to mental health in order to close this gap. Using local interpretable model–agnostic explanations (LIME), the suggested approach produces explanations for model predictions. These explanations are then quantitatively assessed by comparing them to ground-truth explanations using cosine similarity. The models undergo iterative retraining on the enlarged dataset after data samples linked to low-quality explanations are selectively enhanced with text generated by GPT-3.5. The framework is tested on social media–based mental health datasets using a variety of deep learning– and transformer-based models, such as LSTM, Bi-LSTM, BERT variants, and GPT-3.5. According to experimental results, transformer-based models perform better and show steady accuracy gains over time, with overall gains of roughly 4%–7%. The suggested method improves interpretability and predictive accuracy, providing a reliable and scalable solution for high-stakes NLP applications such as reliable mental health classification.
Author supplied keywords
Cite
CITATION STYLE
Das, S., Khondakar, K. R., Mazumdar, H., Kaushik, A., & Singh, S. K. (2026). Enhancing Machine Learning Models for Mental Health Classification Through Iterative Training and Text-Based Augmentation. International Journal of Intelligent Systems, 2026(1). https://doi.org/10.1155/int/2620320
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.