Automated Classification Model for Elementary Mathematics Diagnostic Assessment Data Based on TF-IDF and XGBoost

4Citations
Citations of this article
23Readers
Mendeley users who have this article in their library.

Abstract

With the increasing emphasis on personalized learning, there is a growing demand for automated systems that analyze individual students’ learning states and provide effective feedback. This study proposes a system that analyzes elementary school mathematics diagnostic assessment data to generate personalized feedback. The proposed system integrates Term Frequency-Inverse Document Frequency (TF-IDF) and eXtreme Gradient Boosting (XGBoost) to vectorize textual data and automatically classify learning error patterns. The study utilizes 15,000 diagnostic assessment records collected from 2020 to 2024. After preprocessing, TF-IDF was employed to extract relevant features, and XGBoost was used to train a classification model. To validate the model’s performance, comparative experiments were conducted with Logistic Regression, Support Vector Machine (SVM), LightGBM, BERT, and DistilBERT. The TF-IDF + XGBoost model achieved an accuracy of 98.85% and an F1 Score of 0.9860, outperforming other models. Furthermore, the system demonstrated an average real-time processing speed of 1.3 s, ensuring instant feedback in educational settings. This study highlights the automation and scalability of educational data analysis, suggesting potential applications across various subjects and grade levels.

Cite

CITATION STYLE

APA

Park, S., Oh, S., & Park, W. (2025). Automated Classification Model for Elementary Mathematics Diagnostic Assessment Data Based on TF-IDF and XGBoost. Applied Sciences (Switzerland), 15(7). https://doi.org/10.3390/app15073764

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free