A Comparative Analysis of the Rating of College Students' Essays by ChatGPT versus Human Raters

Potchong M. Jackaria; Bonjovi H. Hajan; Al Rashiff H. Mastul

Journal ArticleOPEN ACCESS

A Comparative Analysis of the Rating of College Students' Essays by ChatGPT versus Human Raters

International Journal of Learning, Teaching and Educational Research (2024) 23(2) 478-492

DOI: 10.26803/ijlter.23.2.23

20Citations

91Readers

Abstract

The use of generative artificial intelligence (AI) in education has engendered mixed reactions due to its ability to generate human-like responses to questions. For education to benefit from this modern technology, there is a need to determine how such capability can be used to improve teaching and learning. Hence, using a comparative-descriptive research design, this study aimed to perform a comparative analysis between Chat Generative Pre-Trained Transformer (ChatGPT) version 3.5 and human raters in scoring students' essays. Twenty essays were used of college students in a professional education course at the Mindanao State University - Tawi-Tawi College of Technology and Oceanography, a public university in southern Philippines. The essays were rated independently by three human raters using a scoring rubric from Carrol and West (1989) as adapted by Tuyen et al. (2019). For the AI ratings, the essays were encoded and inputted into ChatGPT 3.5 using prompts and the rubric. The responses were then screenshotted and recorded along with the human ratings for statistical analysis. Using the intraclass correlation coefficient (ICC), results show that among the human raters, the consistency was good, indicating the reliability of the rubric, while a moderate consistency was found in the ChatGPT 3.5 ratings. Comparison of the human and ChatGPT 3.5 ratings show poor consistency, implying the that the ratings of human raters and ChatGPT 3.5 were not linearly related. The finding implies that teachers should be cautious when using ChatGPT in rating students' written works, suggesting further that using ChatGPT 3.5, in its current version, still needs human assistance to ensure the accuracy of its generated information. Rating of other types of student works using ChatGPT 3.5 or other generative AI tools may be investigated in future research.

Author supplied keywords

Cite

CITATION STYLE

APA

Jackaria, P. M., Hajan, B. H., & Mastul, A. R. H. (2024). A Comparative Analysis of the Rating of College Students’ Essays by ChatGPT versus Human Raters. International Journal of Learning, Teaching and Educational Research, 23(2), 478–492. https://doi.org/10.26803/ijlter.23.2.23

A Comparative Analysis of the Rating of College Students' Essays by ChatGPT versus Human Raters

Abstract

Author supplied keywords

Cite

Register to see more suggestions