Abstract
The limited availability of essay feedback datasets with trait-level annotations has led research in essay evaluation to focus primarily on either automated essay scoring (AES) or feedback generation. In this paper, we introduce LEAF++, an extension of the Learner’s Essay and Feedback (LEAF) dataset, enriched with trait-level scores alongside the original feedback. These trait scores enable a detailed analysis of essay aspects and provide a basis for validating their relationship with feedback. To assess the dataset, we propose four LLM-based validation approaches. Score2Feedback examines whether trait scores guide LLMs to generate feedback similar to the reference feedback. Feedback2Score evaluates whether using reference feedback to revise an essay improves trait scores. FeedbackviaReference iteratively refines feedback, guided by trait scores, until it meets a minimum similarity threshold with the reference. FeedbackviaScore iteratively revises feedback until the revised essay achieves a higher score than the original. These strategies allow us to validate trait scores by measuring their impact on feedback generation and essay revisions, highlighting the relationship between trait-level scores and corresponding feedback. Our experiments using prompting (LLaMA-3, Mistral, Falcon-3, Phi-3) and fine-tuning (LLaMA-3, Mistral, and GPT-OSS) show that trait scores effectively guide LLMs to generate meaningful feedback. The results demonstrate that the annotated trait scores are meaningfully related to the corresponding essay feedback.
Author supplied keywords
Cite
CITATION STYLE
Misgna, H., On, B. W., Lee, I., & Sang Choi, G. (2026). LEAF++: Trait-Annotated Essay Feedback Dataset With Validation of Trait Scores via LLMs. IEEE Access, 14, 9062–9082. https://doi.org/10.1109/ACCESS.2025.3646052
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.