Is Solving Better Than Evaluating GenAI Solutions?

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Generative AI (GenAI) now pervades computer science education, but its pedagogical value depends on how it is integrated. This study explores whether evaluating AI-generated solutions can be as effective as solving problems directly. In a large upper-division algorithms course, we conducted a twelve-week randomized A/B crossover (N = 220) where students alternated between grading ChatGPT’s answers and solving comparable problems themselves. Across exams and overall course grades, performance was statistically indistinguishable between conditions. Homework differences favored whichever cohort encountered the easier half of the syllabus, suggesting task difficulty (as opposed to the evaluation activity) drove those deltas. Surveys showed neutral-to-positive perceptions, with students who reported changing their study habits rating the GenAI-evaluation activity as more helpful. We discuss design choices for GenAI-evaluation tasks that better elicit metacognition without harming achievement.

Cite

CITATION STYLE

APA

Dickey, E., Mertzanidis, M., & Psomas, A. (2026). Is Solving Better Than Evaluating GenAI Solutions? In SIGCSE TS 2026 - Proceedings of the 57th ACM Technical Symposium on Computer Science Education V.2 (pp. 1289–1290). Association for Computing Machinery, Inc. https://doi.org/10.1145/3770761.3777361

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free