Abstract
Large language models (LLMs) are used in high-stakes domains, yet mechanistic interpretability techniques are typically presented as isolated procedures rather than tools supporting human inquiry. We present a human-centered framework that reframes interpretability as a process of sensemaking, experimentation, and judgment through five stages: noticing value tensions, building mental models, interrogating causality, exploring controllability, and making normative judgments. While applicable to any undesirable model behavior, we demonstrate the framework through a case study of demographic bias in LLM-based essay scoring using the PERSUADE 2.0 corpus and Llama-3.1-8B-Instruct. Through linear probing, attribution patching, and activation steering, we show that practitioners can localize where bias emerges and manipulate its effects'achieving up to 83.3% reduction in score gaps while improving accuracy by up to 20.9%. We also uncover a dual mechanism of bias operating through separate pathways, highlighting limits of single-intervention approaches. Our contribution lies in articulating interpretability as an exploratory relationship between humans and models.
Author supplied keywords
Cite
CITATION STYLE
Park, K., Min, H., Lee, C., & Shin, K. Y. (2026). A Human-Centered Framework for Exploring and Steering LLM Internals: A Case Study of Bias in Essay Scoring. In Conference on Human Factors in Computing Systems - Proceedings . Association for Computing Machinery. https://doi.org/10.1145/3772363.3798720
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.