Automated Essay Scoring With GPT-4 for a Local Placement Test: Investigating Prompting Strategies, Intra-Rater Reliability, and Alignment With Human Scores

12Citations
Citations of this article
24Readers
Mendeley users who have this article in their library.
Get full text

Abstract

This study explores the potential of OpenAI's ChatGPT-4 (gpt-4-0613) as an automated essay scoring (AES) tool in a trial involving 300 essays from an American university's academic English program placement test. Three prompting strategies (minimal/detailed rubric, require/not require rationale, and with/without scoring examples) were tested for intra-rater reliability and agreement with human ratings and placements. The study also examined improvements through NLP-derived linguistic features and averaging multiple GPT-4 ratings. Results revealed promising indicators such as high intra-rater reliability, moderate positive correlations with human scores, and better placement agreement than the human average. Prompting strategies influenced scoring accuracy. A prompting strategy combining the provision of a detailed rubric, a request for a scoring rationale, and the provision of scoring examples demonstrated the best performance. Adding linguistic features and averaging multiple GPT-4 ratings were even more effective. The findings suggest that while relying entirely on GPT-4 for essay scoring might be premature, incorporating it as a complement to human ratings could be beneficial.

Cite

CITATION STYLE

APA

Kim, Y. (2025). Automated Essay Scoring With GPT-4 for a Local Placement Test: Investigating Prompting Strategies, Intra-Rater Reliability, and Alignment With Human Scores. TESOL Quarterly, 59(S1), S318–S329. https://doi.org/10.1002/tesq.3405

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free