Human-anchored longitudinal comparison of generative AI with a bias-calibrated LLM-as-judge

0Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Service LLMs evolve without public changelogs, complicating reproducible evaluation. We present a preregistered human-anchored longitudinal study that tracks three major model families over ten weekly waves using a fixed prompt bank (N = 240) across six domains. Blinded human raters provided correctness judgments, and a bias-calibrated LLM-as-judge produced secondary pairwise preferences corrected weekly via a Bradley–Terry model. Mixed-effects modeling and change-point detection (PELT with MBIC penalty) identified significant service drift patterns. Results show divergent stability trajectories among models: one stable, one improving, and one degrading mid-study. Judge calibration increased agreement with humans (τ = 0.59–0.68) while reducing volatility. Safety metrics co-varied with drift events, suggesting behavioral shifts rather than confirmed causal changes. All data, prompts, rubrics, and parameter configurations are provided in supporting files S1–S6.

Cite

CITATION STYLE

APA

Wiese, T. (2026). Human-anchored longitudinal comparison of generative AI with a bias-calibrated LLM-as-judge. PLOS ONE, 21(2 February). https://doi.org/10.1371/journal.pone.0339920

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free