Judging the Judges: A Controlled Audit of Bias and Instability in LLM-as-Judge Evaluation
Main Article Content
Abstract
Large language models are now routinely asked to grade other large language models, a delegation of judgment that has become the default way organizations track progress on open-ended generation tasks where no single reference answer exists. The convenience of this arrangement is undeniable, but it does not remove the underlying measurement problem so much as relocate it inside a system that is itself opaque and variable. This article examines that relocated problem from the position of a practitioner who has had to certify evaluation verdicts as trustworthy enough to justify deployment and governance decisions, not merely observe them from a research distance. Drawing on a growing body of empirical work documenting self-preference effects, superficial-feature rewards, and presentation-order instability in LLM judges, the article synthesizes these findings into a unified bias taxonomy and proposes a controlled matched-pairs audit protocol capable of isolating each bias mechanism while holding response quality constant. Distinct from diagnostic inventories such as EvalBiasBench and quantification frameworks such as CALM, the contribution is a single quality-matched isolation protocol that traces how micro-level verdict shifts propagate into macro-level ranking distortions and that evaluates candidate mitigations on the same matched pairs. We provide a runnable specification, including a pipeline diagram, pair-construction rules, and prompt templates, and we report a small feasibility study on a local judge model (Llama 3.2, 24 matched pairs) in which reversing presentation order flipped 11 of 24 verdicts (45.8%) while substance-free length padding moved only 2 of 14 A-preferring cases to the padded response. The feasibility study is not powered for hypothesis testing; it demonstrates that the instrument runs end to end on a black-box judge and that single-factor isolation is informative even at small scale. The contribution is diagnostic and operational: a reusable framework that reframes LLM judge output as a measurement with a characterizable error profile, appropriate for practitioners who must defend evaluation pipelines to auditors, regulators, or executive stakeholders rather than treat judge scores as ground truth.