Pairwise (Preference) Judging¶
Pairwise judging is a comparison of two responses by a panel of judge models that reconciles their picks into one winner. Instead of asking "is this answer good?" you ask "is A better than B?". It corrects for the position bias that makes a judge favour whichever response it happens to see first.
It is a sibling of LLM as a Jury: same panel machinery, same judge models, but the verdict is a preference (A / B / tie) rather than a pass or a score.
When to use it¶
- You are comparing two systems, prompts, or model versions and want a direct A-vs-B preference rather than two separate absolute scores.
- Absolute grading is hard to calibrate but "which one is better" is clear.
- You want the position bias measured and corrected instead of hoping it washes out.
Quick start¶
llm_jury_pairwise() builds a reusable comparator. Call compare() once per A/B pair:
import asyncio
from evaluatorq import build_report, llm_jury_pairwise
comparator = llm_jury_pairwise(
criteria="The answer is accurate, complete, and directly addresses the question.",
judges=[
"anthropic/claude-sonnet-4-6",
"google/gemini-2.5-pro",
"openai/gpt-5.4-mini",
],
)
async def main() -> None:
comparison = await comparator.compare(
question="What is the capital of France?",
response_a="The capital of France is Paris.",
response_b="The capital of France is Berlin.",
)
print(comparison.winner) # "A"
# Roll many comparisons up into headline and per-judge metrics.
report = build_report([comparison])
print(report.a_win_rate, report.inconclusive_rate)
asyncio.run(main())
llm_jury_pairwise(model="x") is the single-judge shorthand for judges=["x"]. Pass two or more judges to get a panel.
Keep the panel odd and mixed-provider
An odd number of judges (3, 5) makes ties rare. A mix of provider families gives the panel the independence a jury is meant to provide; several judges from the same provider tend to be correlated and add little over one.
Swap and reconcile: how a vote is decided¶
A judge shown response A first and response B second may lean toward the first slot regardless of content. Pairwise judging controls for this by running every judge twice: once as (A, B) and once as (B, A). The second ordering is un-swapped back into the canonical A/B frame, and the two verdicts are reconciled into a single vote:
| First ordering | Second ordering (un-swapped) | Reconciled vote | Flipped |
|---|---|---|---|
A | A | A | no |
tie | tie | tie | no |
A | B | abstains (None) | yes |
A | tie | abstains (None) | yes |
A | missing / failed | abstains (None) | no |
A judge that agrees with itself across both orderings casts that vote. A judge that contradicts itself has no real preference: it flips, abstains from the tally, and the flip is recorded as position bias. The comparison winner is the plurality of the reconciled votes, or inconclusive when no side reaches a plurality or too few judges cast a decisive vote.
Both orderings run concurrently, so swapping does not add wall-clock latency, only cost. Set swap=False to run a single ordering when you have already controlled for position another way; the position-bias metric is then unavailable (no second ordering to disagree with).
Panel configuration¶
llm_jury_pairwise() mirrors llm_jury():
| Argument | Default | What it does |
|---|---|---|
judges | — | Judge model IDs. Two or more makes it a panel. Mutually exclusive with model. |
model | — | Single-judge shorthand for judges=[model]. |
criteria | a general quality rubric | What "better" means. An empty string falls back to the default rubric. |
swap | True | Run both orderings and reconcile. Turn off to skip position-bias correction. |
repetitions | 1 | How many times each judge is asked per ordering; the judge takes its own majority first. |
assignment | "all" | "cyclic" gives each comparison exactly one judge, rotating through the panel (CyclicJudge). Judge bias cancels in expectation across the run at single-judge cost; the assigned judge still runs both orderings when swap is on. Rotation is over the deduplicated panel, repetitions still applies to the one assigned judge, and the cursor lives on the comparator: a reused comparator continues where the previous run stopped. |
replacement_judges | None | Stand-ins for judges that fail mechanically. Promoted per pair and run in both orderings, so a stand-in casts a real reconciled vote. |
min_successful_judges | 1 | Minimum decisive reconciled votes, otherwise the comparison is inconclusive. Must not exceed the panel size. |
state_fields | None | Classify judges only: which template paths are handed to the judge as the material to compare. Defaults to every placeholder the template renders, minus criteria. |
Jev on a pairwise panel¶
typesafe/jev-latest can also sit on a pairwise panel. It receives a three-option choice question, A / B / tie, whose question is your criteria and whose options carry fixed descriptions ("Response A is better", "Response B is better", "Neither is clearly better"). See Jev as a judge for the full classify contract.
comparator = llm_jury_pairwise(
criteria="The answer is accurate, complete, and directly addresses the question.",
judges=["anthropic/claude-sonnet-4-6", "typesafe/jev-latest"],
state_fields=["response_a.output.response", "response_b.output.response"],
)
The material it compares is every placeholder the template renders except criteria itself — for the built-in template that is the question and both responses. Use state_fields to exclude a shared preamble when only the answers matter.
Each ordering rebuilds the question to identify the current A/B seats, then reconciles the verdicts exactly as a prompted judge does.
The literal prompt does not reach a classify judge, but its placeholders still select the default comparison state. system_prompt, temperature, structured_output, max_tokens, reasoning_effort, extra_kwargs and extra_body do not reach it. llm_jury_pairwise() names the non-default settings you set in one warning: when it builds the comparator for known classify ids, or on the first call for a model discovered through the fetched catalogue. They still apply to any prompted judges on the panel.
Because classify verdicts are deterministic, repetitions > 1 warns. Swapping still runs two distinct A/B orderings; repetitions duplicate each ordering and bill for identical answers.
Reading a comparison¶
compare() returns a PairwiseComparison:
comparison.winner # "A" | "B" | "tie" | "inconclusive"
comparison.token_usage # summed across both orderings and any replacements
for vote in comparison.votes:
vote.model # judge model ID
vote.vote # reconciled "A" | "B" | "tie" | None (abstained)
vote.flipped # True if the judge contradicted itself across orderings
vote.completed # True if both orderings were decisive, so a flip was possible
vote.replacement # True if this judge stood in for a failed one
vote.explanation # rationale from the ordering that produced the vote
vote.observations keeps every repetition with its ordering ("ab" or "ba"), zero-based repetition, canonical verdict, and explanation. For a successful classify call, observation.raw_output holds the complete validated answer in the original ordering sent to the provider, including provider-added fields and omitting fields it did not report. A swapped-order choice of "B" therefore remains "B" in raw_output, even when the observation's canonical verdict is "A". Prompted judges and older saved observations have raw_output=None.
Rolling up many comparisons¶
build_report() aggregates a list of comparisons into a PairwiseReport:
report = build_report(comparisons)
report.comparisons # how many went in
report.a_win_rate # A consensus wins over comparisons decided A or B
report.b_win_rate # B consensus wins over comparisons decided A or B
report.tie_rate # consensus ties over all comparisons
report.inconclusive_rate # comparisons the panel could not decide, over all comparisons
report.mean_agreement # mean inter-judge agreement (comparisons with >=2 decisive votes)
for judge in report.per_judge:
judge.model # judge model ID
judge.a_rate # share of its decisive picks that went to A
judge.b_rate # share of its decisive picks that went to B
judge.position_bias # flips over pairs where a flip was possible
judge.tie_rate # ties over all comparisons the judge saw
judge.consistency # shrunk self-consistency from repeated passes, or None
judge.consistency_raw # unshrunk self-consistency, or None
Watch inconclusive_rate alongside the win rates
The win rates are computed over decided comparisons only, so a run that was mostly noise can still show a high a_win_rate. Read it together with inconclusive_rate: a healthy result is a high win rate and a low inconclusive rate. mean_agreement ignores comparisons with a single decisive vote, since one lone voter always "agrees" with itself and would otherwise flatter a degraded panel.
Judge consistency¶
JudgeStats.consistency and JudgeStats.consistency_raw measure whether a judge gives the same answer on repeated passes of the same ordering. They are computed from the recorded observations and are populated under both aggregation="plurality" and aggregation="bt-sigma"; consistency does not require a Bradley-Terry fit.
consistency is the shrunk value used as a reliability weight by repetition-aware BT-sigma, while consistency_raw is the unshrunk self-agreement value. A judge's field is None when that judge has no measurable repeated decisive passes, such as a single-pass run or repeats that all abstained or failed. This means None is an absence of evidence, not a score of zero.
When any judge has consistency data, the HTML report adds Consistency (shrunk) and Consistency (raw) columns under either aggregation. Judges without data show n/a. If no judge has measurable consistency, the report keeps the judge table and adds a caption explaining whether the run used one pass or repeated passes that produced no measurable group.
Reliability-weighted aggregation (BT-sigma)¶
The default consensus treats every judge's vote equally. When your panel mixes judges of different quality (say a frontier model next to a small open-weight one), uniform plurality lets the noisy judges outvote the sharp one. BT-sigma (from "Who can we trust? LLM-as-a-jury for Comparative Assessment", arXiv:2602.16610) fixes this without any labels: it fits a Bradley-Terry model with a per-judge discriminator over the run's own reconciled votes, learning which judges are internally consistent and down-weighting the rest.
from evaluatorq import build_report
report = build_report(comparisons, aggregation="bt-sigma")
report.bt_sigma.p_a_beats_b # fitted global probability that A beats B
report.bt_sigma.judge_sigmas # per-judge discriminator, smaller = more reliable
report.bt_sigma.winners # reliability-weighted winner per comparison
report.bt_sigma.a_win_rate # weighted rollup, next to the plurality one
The headline plurality rates in the report are unchanged, so the two aggregations stay directly comparable, and each JudgeStats entry gains its fitted sigma. The fit is a regularized, unsupervised maximum-likelihood fit on the votes the run already collected: no extra LLM calls or training data. Identical votes are collapsed into weighted counts before the fit (at most three distinct judgements per judge in the A/B setting), so the cost stays flat no matter how many comparisons the run holds. The report exposes fit warnings. When the optimizer does not converge it falls back to uniform plurality only without repetition weights; on a repetition run (below) the winners stay consistency-weighted and only the pooled p_a_beats_b headline degrades to neutral. Do not treat a capped fit as a reliability estimate.
Notes worth knowing:
- Reconciliation already symmetrises position bias (every judge votes in both orderings), which is a requirement of the model, so votes feed the fit as-is.
- With a single judge the discriminator is unidentifiable; the fit falls back to plain Bradley-Terry and says so in
fit_warnings. - A perfectly split panel stays inconclusive rather than letting numerical noise crown one judge reliable.
- A judge whose decisive votes are unanimous (always A, or always B) is excluded from the sigma weighting. With only two items such a judge's sigma measures one-sidedness, not reliability, and
1/sigmawould hand the most degenerate judge on the panel an unbounded weight - the exact shape a position- or verbosity-biased judge takes. On the pooled-fit path it votes with a neutral (median) weight instead; on the repetition path every weight comes from consistency, so that neutral assignment is replaced by the judge's own consistency weight. Either wayfit_warningsnames it. - Check
bt_sigma.converged(andfit_warnings) before trusting sigmas: a fit that stopped at the iteration cap still produces numbers. - Like all unsupervised aggregation, BT-sigma rewards internal consistency. A majority of judges sharing the same systematic bias will still carry the vote; it protects against noisy judges, not coordinated ones.
Repetition-aware reliability¶
Run each comparison more than once (repetitions=2 or more) and repetition-aware BT-sigma can use the consistency values described above as its reliability weights. Instead of the global two-item fit, those weights come from how often each judge agrees with itself on repeated passes of the same prompt.
Two things must be true, and if either is missing the run falls back to the two-item fit and says so in fit_warnings:
- Both orderings must run (
swap=True, the default). Self-agreement is measured inside one ordering, so a single-ordering run (swap=False) never reaches this path, even atrepetitions=2. - At least two judges must have repeated decisive passes. With only one, the fallback weight for every other judge is just that one judge's number, so we do not use it.
Every pass is kept on PairwiseVote.observations. Its verdict is normalized so a swapped-order "B" becomes "A" and the two orderings line up; abstained and failed verdicts are None. Its classifier raw_output stays in the provider's ordering, as described under Reading a comparison.
Why self-agreement, not the global fit. With one pass per judge, the two-item fit cannot tell a noisy judge apart from a hard batch of questions. Repeating the same prompt removes the ambiguity: any disagreement is the judge being inconsistent, nothing else. So a judge's consistency is scored on each prompt, counted once per judge per datapoint, then averaged. Because of that:
- Different questions are never compared against each other, so a hard batch cannot look like an unreliable judge.
- Position bias is not mistaken for inconsistency, since each ordering is its own group.
- Extra repeats sharpen a judge's score but never add weight: each judge still casts one weighted vote per comparison.
Two numbers, and which to read. bt_sigma.repetition_consistency is the reliability weight. It is pulled toward the panel average (so one lucky 2-pass agreement cannot take over the run) and lowered when passes fail, so even a perfectly steady judge reads a little under 1.0 unless the whole panel is steady. bt_sigma.repetition_consistency_raw is the plain self-agreement, with no adjustment: 1.0 means the judge always agreed with itself. The same two values are available as JudgeStats.consistency and JudgeStats.consistency_raw under the report's plurality and BT-sigma aggregations. Read the raw number to compare judges inside one run; do not compare the weight across two runs with different panels, where the same judge can move without changing. p_a_beats_b is still the pooled run-level headline. A judge with no repeats gets the median of the measured weights and is named in fit_warnings. Runs saved before this feature existed still load and behave as before.
What consistency is not. It is self-agreement under fixed conditions, not task difficulty, not judge quality (a judge can be steadily wrong), and not accuracy against a ground truth. A clean abstention does not count against a judge: ['A', <abstention>, 'A'] scores 1.0 in both numbers. A pass that errored or came back off-contract is a failure, not a free abstention, and is counted in repetition_failures. The judge still agreed with itself on the passes it finished, so the raw number for ['A', None, 'A'] stays 1.0; only the weight drops, to 2/3, because a flaky judge is less dependable even when it agrees with itself. The two Nones look identical in the vote list; only repetition_failures tells a clean abstention from a broken pass.
Cost. Calls scale linearly with judges x orderings x repetitions, so R=2 doubles the calls per comparison. Wall-clock barely moves, since the passes run in parallel, and each pass adds one small record. Use R=2 when reliability weighting matters: it is the smallest R that produces any consistency evidence. R=3 only refines the score, for 50% more cost.
For ranking more than two candidates, use evaluatorq.ranking.fit_bt() directly: it takes item pairs from any number of judges and returns skills, a ranking, and per-judge reliability, with cycle_rate() as the matching consistency diagnostic. cycle_rate does not apply to the two-item case above, since two fixed items have no cycles to measure.
Saving a run and viewing it in the dashboard¶
build_report() gives you the numbers in memory. To keep a run and read it in the dashboard, collect the comparisons into a PairwiseRun and save it:
from evaluatorq.pairwise_run import new_run
run = new_run(
run_name="prompt-v2 vs prompt-v3",
label_a="prompt-v2",
label_b="prompt-v3",
judges=["anthropic/claude-sonnet-4-6", "google/gemini-2.5-pro"],
criteria="The answer is accurate, complete, and directly addresses the question.",
)
for question, response_a, response_b in my_pairs:
comparison = await comparator.compare(
question=question, response_a=response_a, response_b=response_b
)
run.add(comparison, question=question, response_a=response_a, response_b=response_b)
run.save() # -> .evaluatorq/pairwise-runs/<timestamp>_prompt-v2-vs-prompt-v3.json
save() rolls the comparisons up with build_report() and stores the result on the run, so the dashboard never recomputes it. Pass a path to choose the file yourself; the default lands in the pairwise run store, where eq dashboard discovers it.
From the project where you saved the run, launch the dashboard with:
With no path, this scans all three default stores: .evaluatorq/runs (red team), .evaluatorq/sim-runs (simulation), and .evaluatorq/pairwise-runs (pairwise). Open the local URL printed by the command.
label_a and label_b name the two systems being compared. They default to "A" and "B", but nothing in the judging data records what was in each slot, so a reader of the dashboard cannot tell what "A won" means. Set them.
A run also records swap. Position bias is only meaningful when both orderings ran, so a run saved with swap=False shows that column as unavailable rather than as a flattering 0.00.
The dashboard renders the run as three sections: the consensus win rates for each side, a per-judge table (win rates, tie rate, position bias, and consistency when measurable), and the comparison list, where each row expands to show the two responses side by side with every judge's vote and rationale.
For classify votes, open Classifier details inside an expanded comparison in the dashboard or standalone HTML report. Each ordering and repetition shows its choice, confidence and probability distribution when reported. The display converts swapped A/B choices and probability keys back to the run's original A/B frame and uses label_a / label_b; tie is unchanged. Missing confidence or probabilities are omitted, never rendered as zero. Percentages near 50% retain enough precision to show which side they fall on and distinguish close values.
Saved PairwiseRun JSON preserves each observation's provider-order raw_output; display conversion does not rewrite it. This is local report data. send_results_to_orq() strips evaluator raw_output, so the hosted Orq experiment view does not receive these classifier details.
The lower-level core¶
run_pairwise() is the ordering-independent engine underneath the comparator. It takes any async judge_fn(first, second, model) rather than building LLM calls itself, so you can drive the swap-and-reconcile logic with your own judge. reconcile_pair() and pairwise_consensus() are exposed for the same reason. Most callers want llm_jury_pairwise(); reach for run_pairwise() when you are plugging in a non-LLM judge or testing the reconciliation directly.
All three live in evaluatorq.pairwise — not evaluatorq.pairwise_run, which holds the run-persistence helpers used above:
run_pairwise() is also re-exported at the top level as evaluatorq.run_pairwise; reconcile_pair() and pairwise_consensus() are not. run_jury() is internal to evaluatorq.common.jury and is not part of the public top-level API.
Judge fan-out is bounded by the shared LLM ceiling: every judge call routes through the shared LLM path and takes a slot, so all judges, orderings, and concurrent compare() calls share one cap. It defaults to 10 concurrent requests; -1 disables that default at the top level. Inside evaluatorq() set it with llm_parallelism=; around a standalone call, wrap it in async with evaluatorq.llm_concurrency_limit(n):. A nested limit can lower the enclosing cap but cannot raise or disable it.
Where to next¶
- LLM as a Jury — score a single response with a judge panel.
- Custom Evaluators & Frameworks — define your own evaluators.
- Red Teaming — the red-team workflow judging plugs into.