Cross-evaluation of Large Language Model Assessment Behaviours in Educational Tasks by Cognitive Level

  • Fedoruk B
N/ACitations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Large language models show promise for educational assessment, but their comparative capability across different cognitive domains remains understudied. This paper presents a systematic analysis of seven leading LLMs—ChatGPT, Claude, Gemini, Perplexity, Mistral, Command R+, and Grok—in their ability to both generate and evaluate educational responses across different levels of the revised Bloom’s taxonomy. Using a novel cross-evaluation methodology, 6,045 evaluations were analyzed using a standard rubric examining content accuracy, cognitive alignment, communication clarity, and response depth. The findings revealed three distinct clusters of grading behaviour: lenient evaluators (Mistral, Gemini, and ChatGPT), moderate evaluators (Claude and Grok), and strict evaluators (Command R+ and Perplexity). Significant variations in grading consistency emerged, with ChatGPT showing the greatest consistency and Perplexity the most variability. Notable systematic biases were observed, including Gemini’s positive bias toward Grok and Command R+’s negative bias toward Gemini. These patterns provide a framework for selecting appropriate LLMs for specific educational tasks while highlighting the importance of understanding their individual evaluation tendencies.

Cite

CITATION STYLE

APA

Fedoruk, B. D. (2025). Cross-evaluation of Large Language Model Assessment Behaviours in Educational Tasks by Cognitive Level. Journal of Educational Informatics, 6(1). https://doi.org/10.51357/jei.v6i1.314

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free