A Human-Centered Framework for Exploring and Steering LLM Internals: A Case Study of Bias in Essay Scoring

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Large language models (LLMs) are used in high-stakes domains, yet mechanistic interpretability techniques are typically presented as isolated procedures rather than tools supporting human inquiry. We present a human-centered framework that reframes interpretability as a process of sensemaking, experimentation, and judgment through five stages: noticing value tensions, building mental models, interrogating causality, exploring controllability, and making normative judgments. While applicable to any undesirable model behavior, we demonstrate the framework through a case study of demographic bias in LLM-based essay scoring using the PERSUADE 2.0 corpus and Llama-3.1-8B-Instruct. Through linear probing, attribution patching, and activation steering, we show that practitioners can localize where bias emerges and manipulate its effects'achieving up to 83.3% reduction in score gaps while improving accuracy by up to 20.9%. We also uncover a dual mechanism of bias operating through separate pathways, highlighting limits of single-intervention approaches. Our contribution lies in articulating interpretability as an exploratory relationship between humans and models.

Cite

CITATION STYLE

APA

Park, K., Min, H., Lee, C., & Shin, K. Y. (2026). A Human-Centered Framework for Exploring and Steering LLM Internals: A Case Study of Bias in Essay Scoring. In Conference on Human Factors in Computing Systems - Proceedings . Association for Computing Machinery. https://doi.org/10.1145/3772363.3798720

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free