Comparing Performance of Feature Extraction Methods and Machine Learning Models in Automatic Essay Scoring

  • Yao L
  • Jiao H
N/ACitations
Citations of this article
11Readers
Mendeley users who have this article in their library.

Abstract

This study used Kaggle data, the ASAP data set, and applied NLP and Bidirectional Encoder Representations from Transformers (BERT) for corpus processing and feature extraction, and applied different machine learning models, both traditional machine-learning classifiers and neural-network-based approaches. Supervised learning models were used for the scoring system, where six out of the eight essay prompts were trained separately and concatenated. Compared with previous study, we found that adding more features such as readability scores using Spacy Textsta improved the prediction results for the essay scoring system. The neural network model, trained on all prompt data and utilizing NLP for corpus processing and feature extraction, performed better than other models with an overall test quadratic weighted kappa (QWK) of 0.9724. It achieved the highest QWK score of 0.859 for prompt 1 and an average QWK of 0.771 across all 6 prompts, making it the best-performing machine learning model that was tested.

Cite

CITATION STYLE

APA

Yao, L., & Jiao, H. (2023). Comparing Performance of Feature Extraction Methods and Machine Learning Models in Automatic Essay Scoring. Chinese/English Journal of Educational Measurement and Evaluation, 4(3). https://doi.org/10.59863/dqiz8440

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free