Est.

Position Bias in LLM Judges Favoring First or Last Speaker

LLM judges systematically favor responses based on position rather than quality alone.

Editorial team · · 10 min read
Cover illustration for “Position Bias in LLM Judges Favoring First or Last Speaker”
Bias Auditing and Fairness · October 5, 2026 · 10 min read · 2,228 words

Position bias in large language model judges is the tendency to favor a response because of where it sits in the sequence presented, not because of anything about its content. This piece explains how that bias forms, why it hides from the usual checks, how it shows up specifically in debate scoring, and what mitigations actually hold up.

What produces position bias

When a language model is asked to compare two responses and pick the better one, the order in which those responses appear can change the outcome on its own, independent of the quality of either response. That is position bias: a structural property of how these models process sequences, not a random error that averages out over enough trials. It appears in two directions. Primacy bias means the judge favors whichever response came first. Recency bias means it favors whichever came last. Neither direction is fixed across all models or all situations. Which way a given judge leans depends on the model family doing the judging, the length of the context window it's working within, and how close the two candidate responses are in actual quality. Close contests push the judge toward positional cues over substantive ones. There's no universal rule that tells you in advance which way a given model will tilt, and a fix that corrects for primacy bias in one model will not automatically correct for recency bias in another. Every new judge has to be tested on its own terms.

Why position bias is harder to detect

The instinct most people have is that a judge that gives the same score when you ask it the same question twice must be a reliable judge. That instinct is wrong, and the gap between consistency and correctness is the center of the problem. A judge can return identical verdicts run after run and still be wrong in exactly the same positional direction every single time, which shows the consistency is evidence of a bias operating on a strict, repeatable schedule. Researchers now call this the consistency-bias paradox. Norman, Rivera, and Hughes, at UC Berkeley, ran what stands as the largest systematic evaluation of LLM-as-a-judge systems to date in June 2026, covering 21 judges from nine different providers across three benchmarks: MT-Bench, JudgeBench, and RewardBench. They found that two production-deployed judges carried test-retest reliability above 0.95, a number that would normally be read as near-perfect consistency, while simultaneously carrying position bias above 0.10, a severe level by the field's own standards. Reliability that high was not a shield against bias that severe. It was sitting right alongside it.

The most common way teams validate these judges, exact-match agreement against human labels, makes the problem worse rather than better, because it overstates how much real agreement exists. The same Norman et al. study found that all 21 judges in the evaluation showed substantial kappa deflation on MT-Bench, with exact match overstating chance-corrected agreement by between 33.8 and 41.3 percentage points. A judge can clear the field's standard validation bar and still carry severe positional distortion underneath that number. Reliability is also not stable across the tests you choose to run it on. Norman et al. found judge rankings shifted by as much as 14 positions depending on which benchmark was used, and the RAND Corporation's Judge Reliability Harness, released by Sunishchal Dev, Morgan Sandler, and colleagues, evaluated four judges across safety, persuasion, misuse, and agentic benchmarks and concluded that none of them held up as reliable across all four. In response, the field's own calibration bar has been getting stricter: the current production target calls for Cohen's kappa above 0.6 against human labels, with anything above 0.8 counted as strong, a bar that a lot of deployed judges passing simple exact-match checks would actually fail.

Diagram: Consistency Is Not the Same as Correctness. Visualizes: Show the paradox that high test-retest reliability and severe position bias coexist in the same judge.

How position bias plays out in debate judging specifically

Apply this to a debate round: whichever speaker happens to go first or last can pick up points from an LLM judge for reasons that have nothing to do with the argument made, making the stakes concrete fast. That's a direct structural problem for any system that hands scoring to a single model judge without any controls around it. Debate rounds have a built-in asymmetry that a general pairwise comparison doesn't carry. The first speaker gets to set the frame the rest of the round is argued against. The last speaker gets the final word before judgment happens. An LLM judge may weight those two positions unevenly without that weighting having anything to do with argument quality, and unlike a generic side-by-side comparison, a debate round can't simply be re-run in reverse order without changing what the format actually is.

The reason this matters goes beyond fairness in any single round. Research on AI safety via debate, from Khan et al. in 2024, found that the debate format helps both non-expert models and human judges answer questions at meaningfully higher accuracy than simpler baseline methods. That gain only holds if the judge evaluating the debate can actually weigh the arguments on their merits. Position bias cuts directly against that condition. The entire case for using debate as a mechanism for getting at truth, rather than just a format for competition, rests on the judge being capable of fair evaluation, and a judge that is quietly rewarding the last speaker regardless of what was argued is not meeting that condition no matter how well-designed the debate format itself is.

The three mitigation strategies that address the problem

Position bias is not something teams have to just accept as the cost of using an LLM judge. Specific structural choices at the system level measurably reduce it, though no single one of them handles the whole problem on its own, and the systems holding up best in practice tend to combine at least two.

The first is order rotation. Run the same comparison twice, once with Response A presented first and once with Response B presented first. If the judge picks A when A goes first and picks B when B goes first, treat the result as a tie rather than letting either ordering decide the outcome. Both Openlayer's 2026 guide and Comet's 2026 guide name this as the primary fix for position bias in pairwise setups, and it's the most direct answer to the mechanism described above: it cancels out the simple preference for a position by forcing the judge to show whether its preference survives a reversal. Rotation has a real limit, though. It corrects for the most direct positional effect, but it doesn't touch other biases that can be operating on the same comparison at the same time, like a preference for longer answers or for answers that sound more authoritative, and a judge with a strong enough underlying positional preference can still end up favoring one ordering even after the results are averaged.

The second is panel judging across multiple model families rather than relying on one model, or several copies of one model, to render the verdict. Different model families carry different positional tendencies, so when several of them are asked to weigh in and their results are aggregated, the panel's combined verdict is less likely to simply reproduce one model's systematic lean. This matters more because of a separate, compounding bias documented by Chae and colleagues at KAIST: LLM judges inflate scores for selections labeled as their own and deflate scores for selections labeled as belonging to another model, and this happens under matched quality conditions, with no stylistic fingerprint present to explain the gap. A panel built entirely from one model family doesn't escape that problem, since same-family preference can still operate across copies of the same underlying model. A 2026 paper by Pombal, Rei, and Martins documented self-preference bias appearing in rubric-based evaluation as well, which is part of why cross-family panel composition is now treated in the field as a core part of how the measurement system is built. It also isn't a settled design choice that works the same way every time: Norman et al. found that two frontier judges, Claude Opus 4.6 and Gemini 3.1 Pro, improved in positional stability moving from MT-Bench to JudgeBench, while judge rankings overall shifted by up to 14 positions across benchmarks in the same study. Current-generation systems behave differently from one another, so panel composition needs to be tested against the specific evaluation task at hand.

The third is a transparent, pre-committed scoring rubric, published before the round begins. A rubric that scores logic, response quality, clarity, and persuasion as separate, named dimensions gives the judge less room to substitute a gut positional read for structured evaluation. Openlayer's 2026 guide points to chain-of-thought prompting, requiring the judge to lay out its reasoning before it delivers a score, as a practice that cuts down on arbitrary decisions and makes flawed logic visible inside the number. Breaking a holistic score into separate criteria correlates better with human judgment than asking the judge for one overall verdict. The order of the rubric itself carries its own bias risk, since criteria listed first in a rubric tend to dominate how the judge weighs the whole evaluation, but publishing a fixed rubric in advance at least makes that effect something that can be checked and audited after the fact. A published rubric also does something beyond constraining the judge in the moment: it gives participants a way to verify, after their round is scored, that the criteria actually applied were the ones that had been announced beforehand.

What the mitigation strategies still don't solve

Combining all three strategies still leaves position bias only reduced, not removed. The realistic goal is making the bias measurable, bounded, and visible, not making it disappear.

Most of the research behind these fixes studies a single judge in isolation, and that leaves a gap once several LLM scorers are combined into a panel: individual biases can interact with each other in a panel in ways that no single-judge fix is built to catch. A panel assembled specifically to cancel out positional preference can still end up amplifying a shared preference for verbosity or for a particular tone of authority if those traits happen to be correlated across the models chosen for the panel. Capability gains in frontier models don't move this problem in one direction either. Norman et al. found that some judges got more positionally stable on harder evaluation items while getting worse on easier ones, meaning model capability and resistance to bias don't move together in a predictable line. Treating "a stronger model" as a stand-in for "a less biased model" is not a safe design assumption.

The sharpest challenge to the standard engineering response is the reliability-validity gap the Norman et al. study reveals. Rotating positions can correct the simplest positional artifact, but the metric most teams use to validate that a judge is working, exact-match agreement, is itself unreliable, so a system can pass that validation standard while still carrying severe positional distortion. It's a reason to validate mitigation against kappa-based calibration measured against human labels, rather than trusting exact-match agreement as a stand-in for whether a judge is actually getting it right. Some of the bias lives at a level that prompt engineering or rubric design cannot reach. Chae et al.'s finding that labels alone, self versus other, shift scores in opposite directions even with no stylistic cue present suggests that part of this bias is built into how these models process identity signals themselves, a layer that sits below anything a better-worded rubric can fix.

Qualities of a judging system that handles position bias

Understanding the mechanism behind position bias and what real mitigation looks like turns a black-box trust problem into something a reader can actually check. Anyone relying on an AI judging system, in debate or anywhere else, should be able to find out whether the scoring criteria were published before the round happened or only produced afterward to explain a result already reached. They should be able to find out whether the system runs each evaluation in more than one ordering or whether the judge only ever sees a single presentation sequence with no reversal check applied. They should be able to find out whether the judges come from more than one model family, since a panel built from a single family can't rule out same-family preference systematically favoring one side. They should be able to see what was actually said cited in the decision itself, so the verdict can be checked against the record rather than handed down as a bare score with no reasoning attached. They should be able to find a real path to appeal a decision, rather than facing a verdict treated as final regardless of whether it can be checked against anything. And they should be able to confirm that the system has been calibrated against human labels using a kappa-based metric, not validated only through exact-match agreement, given how badly that metric has been shown to overstate real agreement. These six questions follow directly from what the research shows works: published rubrics set in advance, order rotation across presentations, cross-family panels, decisions tied to cited evidence, and a real appeals path are not finishing touches on a judging system. They are the mechanisms that separate a judge that merely claims to handle position bias from one that has actually been built to.

Sources

  1. Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
  2. Debating with More Persuasive LLMs Leads to More Truthful Answers

More in Bias Auditing and Fairness