Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments

1Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

A key ethical challenge in Automated Essay Scoring (AES) is ensuring that scores are only released when they meet high reliability standards. Confidence modelling addresses this by assigning a reliability estimate measure, in the form of a confidence score, to each automated score. In this study, we frame confidence estimation as a classification task: predicting whether an AES-generated score correctly places a candidate in the appropriate CEFR level. While this is a binary decision, we leverage the inherent granularity of the scoring domain in two ways. First, we reformulate the task as an n-ary classification problem using score binning. Second, we introduce a set of novel Kernel Weighted Ordinal Categorical Cross Entropy (KWOCCE) loss functions that incorporate the ordinal structure of CEFR labels. Our best-performing model achieves an F1 score of 0.97, and enables the system to release 47% of scores with 100% CEFR agreement and 99% with at least 95% CEFR agreement—compared to ≈ 92% CEFR agreement from the standalone AES model where we release all AM predicted scores.

Cite

CITATION STYLE

APA

Chakravarty, A., Brenchley, M., Breakspear, T., Lewin, I., & Huang, Y. (2025). Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 6, pp. 1498–1507). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.acl-industry.106

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free