Reliable and efficient automated short-answer scoring for a large dataset using active learning and deep learning

3Citations
Citations of this article
17Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

The evaluation of short answers from numerous students requires a highly reliable and efficient scoring system that incorporates natural language processing. Specifically, for higher reliability, higher indicators such as accuracy are desirable, whereas for higher efficiency, less data are desirable for labelling. Moreover, these desires are more acute with larger data sets. In this study, we proposed a novel workflow using active and deep learning to improve the accuracy of automated scoring while reducing the scoring cost incurred through manual scoring. In a trial common test for university entrance examinations, the proposed workflow automatically scored more than 60,000 short answers per question (in Japanese). Consequently, we obtained high accuracy for all the questions in the test. In particular, the proposed workflow achieved an accuracy as high as 0.999, which is consistent with the existing manual grading approach in terms of accuracy. Furthermore, the cost of human scoring was reduced to less than one-sixth that of the training data on average, which is a significant advantage over other machine learning methods. Thus, the proposed workflow demonstrates the potential for advancing practical applications of automated scoring.

Cite

CITATION STYLE

APA

Osaka, J., Maeda, A., Oka, H., Mori, Y., Ishioka, T., & Suyari, H. (2025). Reliable and efficient automated short-answer scoring for a large dataset using active learning and deep learning. Interactive Learning Environments, 33(6), 3776–3787. https://doi.org/10.1080/10494820.2025.2452005

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free