Generative AI in assessing written responses of geography exams: challenges and potential

0Citations
Citations of this article
17Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

This article examines the application of Large Language Models (LLM)–GPT-4, Claude, Cohere, and Llama–to assess students’ open-ended responses in Geography exams. The models’ assessment scores were compared to assessment and scores by the original multi-stage human assessment as well as two additional human expert scoring. The case study considers the high-stakes national matriculation exam in Finland. The exam results play a crucial role in determining individuals’ eligibility for higher education, including a study right in Geography at the university. We selected 18 essays that had originally been given 5 (basic), 10 (good) and 15 (excellent) points on a scale from 0 to 15 points. Findings show variability between LLMs and notable differences between LLM and human evaluations. The language of responses and grading instruction influenced LLM performance. These results highlight the potential and complexities of integrating generative AI today in learning assessments to score open-ended responses. Precise control of prompts and LLM settings proved crucial for the LLM to align with original assessment scores more closely.

Cite

CITATION STYLE

APA

Jauhiainen, J. S., Gagagorry Guerra, A., Nylén, T., & Mäki, S. (2026). Generative AI in assessing written responses of geography exams: challenges and potential. Journal of Geography in Higher Education, 50(2), 210–222. https://doi.org/10.1080/03098265.2025.2593484

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free