LLM as a Jury¶
A jury is a panel of judge models that runs together, aggregates its verdicts into one decision, and reports how much the judges agreed. A single judge can be noisy from one call to the next or biased toward outputs from its own provider family; a jury spreads that decision across several models.
You can use a jury two ways: as a general evaluator in evaluatorq() through llm_jury(), or inside red teaming through EvaluatorConfig. Both share the same panel machinery.
When to use it¶
- The evaluation is high stakes and you want a verdict that does not rest on one model's opinion.
- You are judging outputs from the same provider as your usual judge and want to avoid a judge grading its own family.
- You want a quantitative signal for how much your judges actually agree, so you know when a verdict is solid and when it is contested.
A single judge is cheaper and faster. Reach for a jury when the cost of a wrong verdict outweighs the extra calls. A single-judge panel runs with no aggregation overhead, so a jury is purely additive.
Quick start¶
llm_jury() builds an evaluator you drop into the evaluators=[...] list of evaluatorq(). Give it two or more judges and it becomes a jury:
import asyncio
from evaluatorq import DataPoint, evaluatorq, llm_jury
correctness = llm_jury(
name="correctness",
criteria="The answer is factually correct and directly answers the question.",
judges=[
"anthropic/claude-sonnet-4-6",
"google/gemini-2.5-pro",
"mistral/mistral-large-2411",
],
)
async def answer(data: DataPoint, _row: int) -> dict:
# Your system under test produces the output to be judged.
return {"name": "qa", "output": "Paris is the capital of France."}
async def main() -> None:
await evaluatorq(
"qa-eval",
data=[DataPoint(inputs={"question": "What is the capital of France?"})],
jobs=[answer],
evaluators=[correctness],
)
asyncio.run(main())
llm_jury(model="x") is shorthand for judges=["x"] — a single judge, the classic LLM-as-a-judge. Pass judges=[...] with two or more models to turn it into a jury.
Keep the panel odd and mixed-provider
An odd number of judges (3, 5) makes ties rare. A mix of provider families gives the panel the independence a jury is meant to provide; several judges from the same provider tend to be correlated and add little over one.
Verdict modes¶
verdict_kind (with labels) decides what each judge returns and how passed is set. It is not inferred from labels — pick the mode explicitly.
| Mode | Configure it with | Judge returns | passed is |
|---|---|---|---|
| Boolean (default) | verdict_kind="categorical", no labels | true / false | the boolean itself |
| Labeled | verdict_kind="categorical", labels=[...] | one of labels | verdict in passing_labels (None if passing_labels omitted) |
| Numeric | verdict_kind="numeric" | a float in score_range | score >= threshold |
# Labeled: a fixed rubric, only some labels pass
tone = llm_jury(
name="tone",
criteria="Rate the tone of the reply.",
judges=["anthropic/claude-sonnet-4-6", "google/gemini-2.5-pro"],
labels=["rude", "neutral", "friendly"],
passing_labels=["neutral", "friendly"],
)
# Numeric: a 0-1 score with a pass threshold
helpfulness = llm_jury(
name="helpfulness",
criteria="Score how helpful the answer is, 0 to 1.",
judges=["anthropic/claude-sonnet-4-6", "google/gemini-2.5-pro"],
verdict_kind="numeric",
threshold=0.7,
)
labels/passing_labels are valid only for categorical; passing them with numeric raises ValueError. In labeled mode passing_labels must be a subset of labels; omit it and the verdict is still recorded but passed is None.
Jev as a judge¶
typesafe/jev-latest (Jev) is a classify model: it is never prompted. It is handed the material to judge and one question about it, and answers with a probability distribution instead of prose. On a panel it takes a seat beside prompted LLM judges and casts an ordinary vote.
See Classify judges for the task-focused guide to question shapes, state selection, routing, warnings and failure modes.
A Jev call costs about $0.042 per million input tokens, with free output, and answers in under a second — a three-question probe took 0.7 s. Use it for a cheap third vote on yes/no, fixed-label, or ordered-scale verdicts. Its explanation is synthesised from numbers, so use a prompted judge when you need reasoning.
Seat it by naming it in judges alongside anything else, and give the panel criteria:
from evaluatorq import llm_jury
correctness = llm_jury(
name="correctness",
criteria="The answer is factually correct and directly answers the question.",
judges=["openai/gpt-5.4-mini", "typesafe/jev-latest"],
labels={
"correct": "every claim in the answer is accurate",
"incorrect": "at least one claim is wrong or unsupported",
},
passing_labels=["correct"],
)
labels as a {label: description} dict is what a classify judge needs — the descriptions are the options it picks between. The LLM judge on the same panel gets the same descriptions through its verdict schema and system prompt. labels=["correct", "incorrect"] (a plain list) stays valid and describes nothing.
Numeric mode needs levels instead — 2 to 10 ordered level descriptions, lowest first:
helpfulness = llm_jury(
name="helpfulness",
criteria="How helpful is the answer to the person who asked the question?",
judges=["anthropic/claude-sonnet-4-6", "typesafe/jev-latest"],
verdict_kind="numeric",
levels=["no help at all", "partially helpful", "fully helpful"],
threshold=0.7,
)
The classify path follows these rules:
criteriais the question, and it is required. A known classify id with onlypromptraisesValueErrorat construction. A classify id first discovered through the fetched catalogue fails its first scoring call instead of falling back to prompting. Other panels require at least one ofcriteriaorprompt, and may set both.- The material is the state. A classify judge is handed every placeholder the prompt template renders, minus
criteriaitself (which is the question, not the material).state_fields=["output.response"]narrows it to the paths you name. A path with no value is skipped with a debug log, and a state where nothing at all resolves logs a warning once per evaluator. Values keep their types, soinput.all_messagesremains a message list. - Boolean mode compares against
threshold. The judge answers with a probability and the verdict isprobability >= threshold(default0.5), sothresholdmust lie in[0.0, 1.0]. An LLM judge on the same panel returns a boolean directly and ignores it. - Labeled mode returns the chosen label, exactly as an LLM judge does, so
passing_labelsworks unchanged. - Numeric mode requires
levelsand normalises the raw level-index score asscore / (len(levels) - 1), clamped to[0.0, 1.0]. The resulting verdict always uses that range, soscore_rangehas to stay at its default(0.0, 1.0); an override is rejected. - The explanation is synthesised from the numbers that decided the verdict —
noul=0.92 (threshold 0.5),choice='neutral' (confidence 0.96),score=2.30/4 → 0.57 (confidence 0.81). Values close to a decision boundary retain enough decimal places to show which side they fall on. There is no model-written rationale to report. - Each repetition keeps the complete validated classify answer in
JuryRepetition.raw_output, including confidence and probabilities when reported. It survives in returned and local saved results;send_results_to_orq()stripsraw_output, so the hosted Orq experiment view does not receive it. The judge span also recordsjudge.confidenceandjudge.probabilities(a JSON string) when reported. See Reading the output for the typed path. - Settings a classify judge cannot use are named in a warning. The literal
promptis not sent, but its placeholders still select the default state fields.system_prompt,temperature,structured_output,max_tokens,reasoning_effort,extra_kwargsandextra_bodyhave no role on this path.llm_jury()names the non-default settings you set in one warning: at construction for known classify ids, or on the first call for a model discovered through the fetched catalogue. If another seat may be prompted, the warning says those settings still apply to any prompted judges. repetitions > 1warns. A classify verdict is deterministic, so the extra calls are billed for identical answers.
At runtime, the Orq catalogue's supports_classify flag is authoritative. When a model has no catalogue entry, the built-in fallback covers typesafe/jev-latest and warns once. Before the first call, llm_jury() uses that fallback and any register_model() overrides only for early validation; models discovered through the catalogue receive ignored-setting and repetition warnings on their first call.
A classify judge only works through the Orq router
The classify endpoint lives on the Orq router, so a client that does not route through it never reaches that path. typesafe/jev-latest is then sent as an ordinary chat model, the provider rejects the id, and the panel records a failed judge rather than raising — the symptom is a seat that never votes and a verdict decided by the rest of the panel. Give the run an ORQ_API_KEY (see Configuration), or leave Jev off the panel.
Red teaming and simulation do not seat classify judges. EvaluatorConfig panels are prompted judges only.
Panel configuration¶
| Argument | Default | What it does |
|---|---|---|
preset | None | Name of a ready-made panel from Jury Presets, such as "Balanced Trio". Seats the judges, the aggregation rule and a majority quorum (two of three, three of five). Mutually exclusive with judges and model. |
judges | — | Judge model IDs. Two or more makes it a jury. Mutually exclusive with model. |
model | — | Single-judge shorthand for judges=[model]. With none of preset, judges or model, the panel is one judge on the smart role (openai/gpt-6-sol by default; see Configuration › Models). |
criteria / prompt | — | What the judges go on. Pass at least one; both together is allowed. An LLM judge reads the rendered prompt and can pull the rubric in with a {{criteria}} placeholder; a classify judge is handed criteria as its question and never sees the prompt. Setting both on a panel that seats no classify judge and whose prompt never renders {{criteria}} logs one warning, because nothing then reads the rubric. |
labels | None | Categorical labels, either as a list (["rude", "neutral", "friendly"]) or as a {label: description} dict. The dict form describes each option: an LLM judge reads the descriptions through its verdict schema and system prompt, a classify judge takes them as the options of its choice question. |
levels | None | Numeric mode only: 2 to 10 ordered level descriptions, lowest first. Required when the panel seats a classify judge, whose score comes back scaled into (0.0, 1.0) — so score_range must stay at its default there. |
state_fields | None | Classify judges only: which template paths are handed over as the material to judge. Defaults to every placeholder the prompt template renders, minus criteria. |
repetitions | 1 | How many times each judge is asked. The judge takes its own majority before the panel votes, which smooths per-call noise. |
assignment | "all" | How judges are allocated across datapoints. "all" runs every judge on every datapoint. "cyclic" runs exactly one judge per datapoint, rotating through the panel (see below). |
replacement_judges | None | Stand-in models called only when a configured judge fails mechanically. |
min_successful_judges | None | Minimum decisive judges required, otherwise the verdict is inconclusive. Must not exceed the panel size. None means 1 for a hand-listed panel, and a majority of the seats under a preset (two of three, three of five). It is a floor on how many judges must answer, not the threshold the aggregator applies. |
threshold | 0.5 | Numeric mode: passed when score >= threshold. In boolean mode it is the cutoff a classify judge's probability must clear to read as true; an LLM judge returns a boolean and ignores it. |
structured_output | True | Use the provider's structured-output API; falls back to a schema-injected json_object call for models that reject it. |
Cyclic assignment (CyclicJudge)¶
assignment="cyclic" gives each datapoint exactly one judge, rotating through the panel so every judge covers an equal share of the run. Judge bias still cancels in expectation across the dataset, but the run costs the same as a single-judge evaluation instead of len(panel) times as much.
jury = llm_jury(
name="quality",
criteria="Is the answer helpful and correct?",
judges=["openai/gpt-5.4-mini", "openai/gpt-5.4-nano", "deepseek/deepseek-v4-flash"],
assignment="cyclic",
)
Use it for run-level scores (a benchmark mean, a pass rate); keep the default "all" when an individual verdict has to stand on its own, since each per-item verdict under "cyclic" is one judge's opinion and stats/raw_agreement come back None.
See Cyclic judge assignment for how items map to judges, auditing the rotation via raw_output["jury"], and the failure semantics.
Routers as a judge¶
A single judge can be an orq/* router (see Routers). Do not build a panel out of them.
A router optimises each request on its own, so a panel of three routers can return three cards from one vendor, and a panel whose judges share a lineage shares their blind spots too. Cancelling correlated error is the whole reason to poll three judges rather than ask one judge three times, so a jury names its models: either a preset, which seats one model per lineage deliberately, or your own judges list.
The same caution applies to a single-judge run you intend to repeat. A router picks fresh every call, so a re-run can be scored by a different judge than the first pass, and a scoring difference then has two possible causes instead of one.
How the verdict is decided¶
- Each judge votes. With
repetitions > 1a judge is asked several times and reduces its own passes to one vote first (plurality for categorical, mean or median for numeric). - Failures pull in replacements. For every configured judge that fails mechanically, one model from
replacement_judgesstands in, up to the number of failures. - The panel aggregates. Categorical verdicts are decided by plurality vote; numeric verdicts by mean or median.
- Thresholds and ties apply. If fewer than
min_successful_judgesreturn a usable verdict, the result is inconclusive.
A judge can also abstain: it returns cleanly but declines to choose. An abstention is not a failure and does not trigger a replacement, but it is excluded from the decisive tally.
Reading the output¶
llm_jury() returns a standard evaluator, so each result carries the aggregated verdict in value, the pass/fail in passed, and a human-readable panel breakdown (who voted what, how close it was) appended to explanation:
results = await evaluatorq(..., evaluators=[correctness])
for r in results:
for job in r.job_results:
for score in job.evaluator_scores:
print(score.evaluator_name, score.score.value, score.score.pass_)
print(score.score.explanation) # includes the per-judge jury summary
The appended summary is one line ([jury: 3/3 judges, raw agreement 100%]) — it does not name the individual judges or say what each of them decided. That detail rides on raw_output["jury"], which every judged result carries: the serialized JuryResult with one entry per judge (model ID, verdict, explanation, abstain/failure state) plus panel stats and raw agreement. When repetitions > 1, each vote also carries that judge's raw per-repetition verdicts and the reasoning behind each one.
The payload is the JSON form of JuryResult, so validate it back into the typed model rather than indexing raw keys — that is what the red-team report layer does with the same payload (redteam/reports/converters.py):
from evaluatorq.contracts import JURY_RAW_OUTPUT_KEY, JuryResult
jury = JuryResult.model_validate(score.score.raw_output[JURY_RAW_OUTPUT_KEY])
for vote in jury.votes:
print(vote.model, vote.value, vote.explanation)
for rep in vote.repetitions:
print(" ", rep.value, rep.explanation, rep.raw_output)
print(jury.raw_agreement, jury.stats)
Each entry in vote.repetitions is a JuryRepetition with a value (None where the pass abstained or failed), an explanation (None where the pass produced no text), and optional raw_output. A successful classify pass stores its complete validated answer in raw_output, including provider-added fields and omitting fields the provider did not report. Prompted judges leave it None. Each repetition retains its own dictionary; distributions are not averaged into the final verdict. Older reports without raw_output still load with None; legacy bare verdicts also load with explanation=None.
send_results_to_orq() strips raw_output before upload, so the hosted Orq experiment view does not receive these details. They remain available in returned results and local saved result artifacts.
A datapoint whose target errored is never judged: it returns inconclusive with raw_output still None, because no panel ran. Reach for .get(JURY_RAW_OUTPUT_KEY) if your code walks every result rather than only the judged ones.
When the panel did run but reached no verdict, the result is inconclusive and raw_output carries a second key, EVAL_ERROR_RAW_OUTPUT_KEY ("evaluation_error"), naming why: the last judge error, the full list of judge errors, and how many judges failed. It is written only when a judge actually recorded a failure — a panel where every judge abstained cleanly, or where too few of them agreed, produced a non-verdict rather than an outage and gets no error payload. A single shared writer keeps this key beside "jury"; in a direct llm_jury() result, the payload stays in raw_output for your code to read, while the red-team report converter additionally promotes it to result.evaluation_error for the run's error rollup.
The judge span separately records reported classify confidence and probabilities; stripping the experiment upload does not remove those trace attributes.
In red teaming¶
Red teaming reaches the same panel through EvaluatorConfig, where the verdict is the categorical RESISTANT/VULNERABLE case (passed=True means RESISTANT):
from evaluatorq.redteam import EvaluatorConfig, LLMConfig, OpenAIModelTarget, red_team
report = await red_team(
target=OpenAIModelTarget(model="gpt-5.6-luna"),
llm_config=LLMConfig(
evaluator=EvaluatorConfig(
judges=[
"anthropic/claude-sonnet-4-6",
"google/gemini-2.5-pro",
"mistral/mistral-large-2411",
],
min_successful_judges=2,
strict_panel=True, # refuse a judge that shares the target's family
),
),
mode="dynamic",
categories=["LLM01"],
max_dynamic_datapoints=3,
)
strict_panel only fires on a same-family judge
strict_panel=True raises ValueError when any judge shares the target's provider family — same-family self-judging can bias the verdict toward the target's own provider. The panel above is entirely cross-family against an OpenAI target, so the guard passes silently (the healthy case). It would raise only if you added an in-family judge such as "openai/gpt-5.6-luna".
min_successful_judges vs. min_evaluation_coverage — two different levels
min_successful_judges above is a per-attack quorum: it decides whether this one jury panel produced enough decisive votes to reach a verdict for this one attack. EvaluatorConfig also has min_evaluation_coverage (default 0.8), which is a separate, run-level floor: the fraction of all attacks in the run that must get any verdict at all. A min_successful_judges miss on one attack is exactly what produces one of the unevaluated attacks that min_evaluation_coverage counts against. Missing the run-level floor makes eq redteam run exit 1 — see Red Teaming › In CI.
EvaluatorConfig adds strict_panel (turn panel-composition warnings into hard errors) and surfaces a per-attack jury breakdown plus a run-level reliability statistic:
for result in report.results:
jury = result.evaluation.jury if result.evaluation else None
if jury is None:
continue
print(f"{jury.judges_succeeded}/{jury.judges_configured} judges, agreement {jury.raw_agreement}")
for vote in jury.votes:
print(vote.model, vote.value, vote.abstained, vote.error)
reliability = report.summary.jury_reliability
if reliability:
print(reliability.krippendorff_alpha) # 1.0 = perfect, ~0 = chance, <0 = systematic disagreement
Reliability, in short¶
For red-team runs, raw_agreement tells you how lopsided one vote was, and Krippendorff's alpha on the run tells you whether your judges agree more than they would by chance:
1.0is perfect agreement.- around
0is chance level, so the panel is not adding signal. - below
0is systematic disagreement, which usually means the judges are reading the rubric differently and the prompt or panel needs another look.
It is None when undefined, for example a single-judge run or fewer than two multi-judge samples.
Endpoint and retries¶
llm_jury() and PairwiseComparator always send judge calls to the Orq router's responses endpoint (api='responses') — the endpoint the router prices, so a jury verdict records cost the same way the call it judges does. There is no config knob to opt out; run_judge handles the cases that cannot use it on its own, falling back to Chat Completions for a client that does not route through the Orq router, a model the catalogue does not qualify for Responses, or a model that 400s on the endpoint.
A classify judge is the exception: it serves neither endpoint, so api='responses' says nothing about it and it goes to the router's classify endpoint instead. Which of the three served a verdict is recorded on JudgeOutcome.endpoint.
Retries happen at the run_judge layer, not the SDK's: with a default budget of one retry (LLMCallConfig.retry_count=1, i.e. up to two requests per judge call), run_judge disarms the SDK-level retry budget (max_retries=0) on whichever client it is given — the one either helper builds for itself when no client= is passed in, or a client you pass in yourself — so the two retry layers never stack. There is no way to keep an injected client's own SDK retries active alongside run_judge's.
Reasoning effort on a jury¶
llm_jury(), llm_jury_pairwise() and PairwiseComparator take a reasoning_effort= argument that pins the effort on the judge model:
from evaluatorq import llm_jury
jury = llm_jury(
name='helpfulness',
criteria='Is the answer helpful and correct?',
judges=['openai/gpt-5.6-luna', 'anthropic/claude-sonnet-5'],
reasoning_effort='high',
)
Because these judges send to the responses endpoint, the effort renders as a reasoning block — unless run_judge falls back to Chat Completions for one of the reasons above, in which case it renders as a flat reasoning_effort field. Either way the provider is the authority on which values it accepts. A classify judge on the panel ignores the effort and says so in a warning — it takes no sampling settings at all.
This is a distinct knob from red teaming's target_reasoning_effort (the agent under test), from LLMCallConfig.reasoning_effort on red teaming's own attacker= / evaluator= roles, and from EVALUATORQ_REASONING_EFFORT (the simulator's user-simulator and judge). See Tuning for the full disambiguation.
Provider options¶
llm_jury(), llm_jury_pairwise() and PairwiseComparator take both injection seams, and they are not interchangeable:
jury = llm_jury(
name='helpfulness',
criteria='Is the answer helpful?',
judges=['openai/gpt-5.6-luna'],
extra_kwargs={'top_p': 0.9}, # top-level call argument
extra_body={'my_router_field': 'x'}, # request body field
)
extra_kwargs sets arguments on the SDK call and replaces the key. extra_body adds fields to the request body and is merged per key, so router-owned body fields survive alongside yours. Both reach prompted judges only — a classify judge ignores them and names them in a warning.
extra_body inside extra_kwargs is rejected
extra_body is one of the structural fields the call site owns, so passing it through extra_kwargs raises — and because a judge failure is caught and turned into a verdict, the symptom is a judge that always fails rather than a crash. Use the extra_body= parameter.
Full example¶
A complete red-teaming script covering repetitions, replacements, the min_successful_judges threshold, strict_panel, and reading the per-result and run-level output lives at examples/redteam/16_llm_as_a_jury.py.
Where to next¶
- Jury Presets: named panels with published costs, instead of picking judges yourself.
- Pairwise Judging — compare two responses instead of scoring one.
- Custom Evaluators & Frameworks — define your own evaluators.
- Red Teaming — use a jury as the red-team verdict.