Abstract
We fine-tuned and compared several encoder-based Transformer large language models (LLM) to predict differential item functioning (DIF) from the item text. We then applied explainable artificial intelligence (XAI) methods to identify specific words associated with the DIF prediction. The data included 42,180 items designed for English language arts and mathematics summative state assessments among students in grades 3 to 11. Prediction (Formula presented.) ranged from.04 to.32 among eight focal and reference group pairs. Our findings suggest that many words associated with DIF reflect minor subdomains included in the test blueprint by design, rather than construct-irrelevant content that may need to be removed from assessments. This may explain why qualitative reviews of DIF items often yield inconclusive results. Our approach can be used to (1) screen words associated with DIF during the item-writing process for immediate revision to reduce preventable adverse DIF, (2) assist traditional DIF item reviews by highlighting key words, or (3) use DIF prediction as an alternative when obtaining sufficient sample size for traditional DIF analyses is impossible. Extensions of this research can enhance the assessment fairness, especially programs that lack resources to build high-quality items, and among smaller subpopulations with insufficient sample sizes for traditional DIF analyses. See source code here.
Cite
CITATION STYLE
Maeda, H., & Lu, Y. (2025). Finding Words Associated with DIF: Predicting Differential Item Functioning Using LLMs and Explainable AI. Journal of Educational Measurement, 62(4), 883–906. https://doi.org/10.1111/jedm.70017
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.