evaluatorq¶
evaluatorq async ¶
evaluatorq(
name: str,
params: EvaluatorParams | dict[str, Any] | None = None,
*,
data: DatasetIdInput
| ExperimentInput
| TraceInput
| Sequence[Awaitable[DataPoint] | DataPointInput]
| None = None,
jobs: list[Job] | None = None,
evaluators: list[Evaluator] | None = None,
datapoint_parallelism: int | None = None,
llm_parallelism: int | None = None,
parallelism: int | None = None,
print_results: bool = True,
description: str | None = None,
path: str | None = None,
inference: bool | None = None,
single_trace: bool = False,
on_datapoint_complete: DataPointComplete | None = None,
) -> EvaluatorqResult
Run an evaluation with the given parameters.
Can be called with either a params dict/object or keyword arguments:
# Using keyword arguments (recommended):
await evaluatorq("name", data=[...], jobs=[...], datapoint_parallelism=5)
# Using a dict:
await evaluatorq("name", {"data": [...], "jobs": [...], "datapoint_parallelism": 5})
# Using EvaluatorParams:
await evaluatorq("name", EvaluatorParams(data=[...], jobs=[...]))
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name | str | Name of the evaluation run | required |
params | EvaluatorParams | dict[str, Any] | None | Optional EvaluatorParams instance or dict with all parameters. | None |
data | DatasetIdInput | ExperimentInput | TraceInput | Sequence[Awaitable[DataPoint] | DataPointInput] | None | The data to evaluate. A DatasetIdInput to fetch from Orq platform, an ExperimentInput to replay an experiment's recorded responses (requires inference=False), a TraceInput to import recorded trace conversations (requires inference=False), or a list of DataPoint instances/awaitables. | None |
jobs | list[Job] | None | The jobs to run on the data. | None |
evaluators | list[Evaluator] | None | The evaluators to use. If not provided, only jobs will run. | None |
datapoint_parallelism | int | None | Task concurrency, applied at two levels: datapoints run at most | None |
parallelism | int | None | Deprecated alias for | None |
llm_parallelism | int | None | Ceiling on in-flight LLM requests for the whole run, counted per request rather than per task, so it holds however the datapoint/job/evaluator/jury fan-out nests. Defaults to 10; -1 disables it. This is the knob to set against a provider concurrency limit; | None |
print_results | bool | Whether to print results table to console. Defaults to True. | True |
description | str | None | Optional description for the evaluation run. | None |
path | str | None | Optional path (e.g. "MyProject/MyFolder") to place the experiment in a specific project and folder on the Orq platform. | None |
inference | bool | None | Leave unset to resolve it from | None |
single_trace | bool | Group every row under one | False |
on_datapoint_complete | DataPointComplete | None | Optional sync or async callback invoked exactly once after each DataPointResult reaches a terminal state. Exceptions propagate and abort the run. | None |
Returns:
| Type | Description |
|---|---|
EvaluatorqResult | List of DataPointResult objects |
Raises:
| Type | Description |
|---|---|
ValidationError | If parameters fail validation. |
ValueError | If neither params nor required kwargs are provided. |
Example
from evaluatorq import DataPoint, EvaluationResult, evaluatorq, job
@job("uppercase")
async def uppercase_job(data: DataPoint, row: int):
return data.inputs["text"].upper()
async def matches_expected(params):
return EvaluationResult(value=1 if params["output"] == params["data"].expected_output else 0)
await evaluatorq(
"uppercase-eval",
data=[DataPoint(inputs={"text": "hi"}, expected_output="HI")],
jobs=[uppercase_job],
evaluators=[{"name": "matches-expected", "scorer": matches_expected}],
)
deployment async ¶
deployment(
key: str,
inputs: dict[str, object] | None = None,
context: dict[str, object] | None = None,
metadata: dict[str, object] | None = None,
thread: ThreadConfig | None = None,
messages: list[MessageDict] | None = None,
) -> DeploymentResponse
Invoke an Orq deployment and return the response content.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
key | str | The deployment key (name) | required |
inputs | dict[str, object] | None | Input variables for the deployment template | None |
context | dict[str, object] | None | Context attributes for routing | None |
metadata | dict[str, object] | None | Metadata to attach to the request | None |
thread | ThreadConfig | None | Thread configuration for conversation tracking. Must include 'id' key. | None |
messages | list[MessageDict] | None | Chat messages for conversational deployments | None |
Returns:
| Type | Description |
|---|---|
DeploymentResponse | DeploymentResponse with content and raw response |
Example
# Simple invocation
response = await deployment("my-deployment")
print(response.content)
# With inputs
response = await deployment("summarizer", inputs={"text": "Long text..."})
# With messages for chat-style deployments
response = await deployment("chatbot", messages=[
{"role": "user", "content": "Hello!"}
])
# With thread tracking
response = await deployment("assistant",
inputs={"query": "What is AI?"},
thread={"id": "conversation-123"}
)
llm_jury ¶
llm_jury(
*,
name: str,
criteria: str | None = None,
prompt: str | None = None,
system_prompt: str | None = None,
preset: str | None = None,
judges: list[str] | None = None,
model: str | None = None,
repetitions: int = 1,
assignment: Literal['all', 'cyclic'] = 'all',
replacement_judges: list[str] | None = None,
min_successful_judges: int | None = None,
verdict_kind: Literal['categorical', 'numeric'] = 'categorical',
labels: list[str] | dict[str, str | None] | None = None,
levels: list[str] | None = None,
passing_labels: list[str] | None = None,
state_fields: list[str] | None = None,
aggregator: AggregatorSpec | None = None,
threshold: float = 0.5,
score_range: tuple[float, float] = (0.0, 1.0),
tie_break: TieBreak | None = None,
structured_output: bool = True,
temperature: float | None = None,
max_tokens: int = DEFAULT_JURY_MAX_TOKENS,
timeout_ms: int = 90000,
extra_kwargs: dict[str, Any] | None = None,
extra_body: dict[str, Any] | None = None,
reasoning_effort: str | None = None,
client: Any = None,
) -> Evaluator
Build a jury (or single-judge) LLM evaluator for evaluators=[...].
Presets¶
preset seats a named panel from evaluatorq.jury_presets instead of listing judges by hand — llm_jury(name='helpfulness', preset='Balanced Trio'). The seats are derived from the committed model-garden snapshot and carry a published $/1k that a test recomputes, so a preset cannot quietly become a different or a differently-priced panel.
It fills in three things and contradicts none: the judges (so judges/ model alongside it is an error, not an override), aggregator ("majority", or the preset's numeric rule) and min_successful_judges (a majority of the seats, two of three or three of five, which is a floor on how many judges must answer rather than the threshold the aggregator then applies). Pass either of the last two explicitly to overrule it.
Presets are pointwise panels, so assignment="cyclic" is rejected: a rotation runs one judge per item and there is no panel left to agree.
What a preset cannot express yet is per-judge call settings. Its seats run at their catalogue default reasoning efforts — JuryPreset.seated_efforts() reports them — and reasoning_effort here overrides all of them at once (RES-1347).
Provider options¶
extra_kwargs and extra_body are the two injection seams, and they are not interchangeable. extra_kwargs sets top-level arguments on the SDK call and replaces the key; it rejects the structural fields the call site owns (model, input/messages, text/response_format, extra_body). extra_body adds fields to the request body and is merged per key, so router-owned body fields survive alongside yours.
Reach for extra_body for anything the provider reads out of the body that the SDK has no named parameter for. Passing it inside extra_kwargs raises at judge time, not at construction.
Criteria and prompt¶
Pass criteria, prompt, or both — at least one, because a judge with neither has nothing to go on. An LLM judge reads the rendered prompt, which can pull the rubric in with a {{criteria}} placeholder; a classify judge is handed criteria directly as its question and never sees the prompt. Setting both on a panel that seats no classify judge and whose prompt never renders {{criteria}} logs one warning: nothing would read the rubric.
Classify judges¶
A classify model (typesafe/jev-latest) is not prompted. It is handed the material to judge and a question about it, so criteria is required on a panel that seats one, and prompt, system_prompt and the sampling settings play no part in its verdict. These four keywords shape that question, and three of them reach an LLM judge on the same panel too:
| keyword | what it does |
|---|---|
labels={'good': 'fully correct', 'bad': None} | the dict form of labels: the options of a choice question, with what each one means. An LLM judge gets the same descriptions through the verdict schema and the default system prompt. labels=[...] stays valid and describes nothing |
levels=['useless', 'ok', 'excellent'] | 2-10 ordered level descriptions for a score question. verdict_kind="numeric" only, required there when a classify judge is seated, and score_range must stay (0.0, 1.0) because the score comes back scaled into it |
state_fields=['output.response'] | which template paths are handed over as the material to judge. Defaults to every placeholder the prompt template renders, minus criteria (which is the question, not the material). A path with no value is skipped |
threshold | in boolean mode, the cutoff the classify judge's probability must clear to read as true. An LLM judge returns a boolean and ignores it (in numeric mode it is the pass/fail cutoff for both) |
Verdict modes¶
The judge's verdict type and the passed (pass/fail) field are decided by verdict_kind together with labels. verdict_kind is not inferred from labels — it defaults to "categorical" and you pick the mode explicitly. There are three modes:
| how you configure it | judge returns | passed is |
|---|---|---|
boolean mode — verdict_kind="categorical" (default), labels=None | a JSON boolean true/false | the boolean itself |
labeled mode — verdict_kind="categorical" with labels=[...] | one of labels (a string) | verdict in passing_labels (None if no passing_labels given) |
numeric mode — verdict_kind="numeric" | a float in score_range | score >= threshold |
Notes
- For a yes/no judge, use boolean mode (the default — just omit
labels).passedis populated automatically; you do not needpassing_labels. labelsandpassing_labelsare valid only forverdict_kind="categorical"; passing them with"numeric"raisesValueError.- In labeled mode,
passing_labelsmust be a subset oflabels. If you omit it, the verdict is still recorded butpassedisNone(no pass/fail, so no pass-rate to aggregate).passing_labelsrequireslabels— it is rejected in boolean mode (which derives pass/fail from the verdict directly). labelsmust be strings. NativeTrue/Falseare not labels — use boolean mode for that.-
aggregatorpicks the panel consensus rule and must match the verdict kind (a mismatch raisesValueError): -
categorical:
"mode"(default — most common; plurality ties go totie_break) or"majority"(strict >50%, else inconclusive). - numeric:
"mean_std"(default — mean verdict; std reported instatson a conclusive verdict),"median","min", or"max". - a custom
Callable[[list[JuryVote]], bool | float | str | None]for either kind (returnNonefor "no consensus" / inconclusive). The same numeric keyword also collapses a single judge'srepetitions. -
assignmentpicks how judges are allocated across datapoints: -
"all"(default): every judge scores every datapoint — per-item consensus at K times single-judge cost. "cyclic": round-robin (CyclicJudge, arXiv:2603.01865) — each datapoint is scored by exactly one judge, cycling through the panel so every judge covers an equal share. Panel-relative judge bias cancels in expectation over the run at single-judge cost; per-item verdicts are single-judge opinions, so use it for benchmark/run-level scores, not when each individual verdict must be trustworthy. Per-itemstatsandraw_agreementareNone: one vote has no cross-judge agreement to report. The rotation runs over the deduplicated panel, so listing a judge twice to up-weight it only works under"all".repetitionsstill applies to the single assigned judge — N calls to one judge, never N judges. Insideevaluatorq()the assignment is keyed on the dataset row (datapointigoes to judgei % len(panel)), deterministic at anyparallelismand across evaluator reuse. A direct scorer call carries no row and rotates in arrival order on a cursor that lives on the evaluator — only the equal share balance is guaranteed there. Shuffle the dataset first if its order is meaningful. Requiresmin_successful_judges=1; a failed item degrades to inconclusive unlessreplacement_judgesis set.- Every judged result carries the full
JuryResultonEvaluationResult.raw_outputunder the"jury"key: per-judge votes with each judge's model ID, verdict, explanation, and raw per-repetition verdicts and reasoning, plus panel stats and raw agreement.explanationonly gets a one-line summary, so this is the only place the per-judge detail survives. A datapoint whose target errored is never judged and returnsinconclusivewithraw_output=None— no panel ran, so there is no record to carry.
Examples¶
Boolean — "is the answer correct?" (verdict_kind="categorical", the default):
Labeled categorical (verdict_kind="categorical"):
llm_jury(
name="grade",
criteria="Grade the answer.",
labels=["correct", "partially_correct", "incorrect"],
passing_labels=["correct", "partially_correct"],
)
Numeric (verdict_kind="numeric"):
EvaluatorQ Python - An evaluation framework for LLM applications.
job ¶
job(
name: str, fn: Callable[[DataPoint, int], Awaitable[Output] | Output] | None = None
) -> Job | Callable[[Callable[[DataPoint, int], Awaitable[Output] | Output]], Job]
Helper function/decorator to create a named job that ensures the job name is preserved even when errors occur during execution.
This wrapper:
- Automatically formats the return value as {"name": ..., "output": ...}
- Attaches the job name to errors for better error tracking
- Can be used as a decorator (@job("name")) or function (job("name", fn))
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name | str | The name of the job | required |
fn | Callable[[DataPoint, int], Awaitable[Output] | Output] | None | The job function that returns the output (optional when used as decorator) | None |
Returns:
| Type | Description |
|---|---|
Job | Callable[[Callable[[DataPoint, int], Awaitable[Output] | Output]], Job] | A Job function that always includes the job name |
Example
# As a decorator:
@job("text-analyzer")
async def analyze_text(data: DataPoint, row: int):
return {"length": len(data.inputs["text"])}
# As a function wrapper:
my_job = job("my-job", async_function)
# With lambda for simple cases:
uppercase_job = job("uppercase", lambda data, row: data.inputs["text"].upper())
DataPoint ¶
Bases: BaseModel
A data point for evaluation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
inputs | The inputs to pass to the job. | required | |
expected_output | The expected output of the data point. Used for evaluation and comparing the output of the job. | required |
Example
Evaluator ¶
EvaluationResult ¶
EvaluationResultCell ¶
Bases: BaseModel
string_contains_evaluator ¶
string_contains_evaluator(
*, case_insensitive: bool = True, name: str = 'string-contains'
) -> Evaluator
Creates an evaluator that checks if the output contains the expected output. Uses the data.expected_output from the dataset to compare against.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
case_insensitive | bool | Whether the comparison should be case-insensitive | True |
name | str | Optional name for the evaluator | 'string-contains' |
Returns:
| Type | Description |
|---|---|
Evaluator | An Evaluator that checks if output contains expected output |
DatasetIdInput ¶
Bases: BaseModel
Input for fetching a dataset from Orq platform.
invoke async ¶
invoke(
key: str,
inputs: dict[str, object] | None = None,
context: dict[str, object] | None = None,
metadata: dict[str, object] | None = None,
thread: ThreadConfig | None = None,
messages: list[MessageDict] | None = None,
) -> str
Invoke an Orq deployment and return just the text content. This is a convenience wrapper around deployment() for simple use cases.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
key | str | The deployment key (name) | required |
inputs | dict[str, object] | None | Input variables for the deployment template | None |
context | dict[str, object] | None | Context attributes for routing | None |
metadata | dict[str, object] | None | Metadata to attach to the request | None |
thread | ThreadConfig | None | Thread configuration for conversation tracking. Must include 'id' key. | None |
messages | list[MessageDict] | None | Chat messages for conversational deployments | None |
Returns:
| Type | Description |
|---|---|
str | The text content of the response |
llm_jury_pairwise ¶
llm_jury_pairwise(
*,
judges: list[str] | None = None,
model: str | None = None,
criteria: str | None = None,
prompt: str | None = None,
system_prompt: str | None = None,
swap: bool = True,
repetitions: int = 1,
assignment: Literal['all', 'cyclic'] = 'all',
replacement_judges: list[str] | None = None,
min_successful_judges: int = 1,
max_tokens: int = DEFAULT_JURY_MAX_TOKENS,
timeout_ms: int = 90000,
temperature: float | None = None,
structured_output: bool = True,
extra_kwargs: dict[str, Any] | None = None,
extra_body: dict[str, Any] | None = None,
reasoning_effort: str | None = None,
client: Any = None,
state_fields: list[str] | None = None,
) -> PairwiseComparator
Build a pairwise (A-vs-B) LLM jury that reuses the shared panel machinery.
Judges compare two responses and pick a winner ('A'/'B'/'tie'). With swap on (default) each judge is run in both orderings to correct for position bias (see ADR-24). Panel/orchestration params mirror llm_jury. prompt overrides the built-in Mustache-style template (which exposes the response_a.*/response_b.* namespace via _side_to_namespace); leave it None to use the default. Returns a PairwiseComparator; call compare per A/B pair, and roll many comparisons up with evaluatorq.pairwise.build_report.
A classify judge (typesafe/jev-latest) is seated the same way it is on a pointwise panel: it answers an A/B/tie choice question whose question is criteria and whose material is every placeholder the template renders except criteria itself. state_fields narrows that material to the paths you name. Position-bias swapping is unchanged — the question is rebuilt per ordering, so each call says which response sits in which seat.
Concurrency is bounded by the run-scoped LLM ceiling (see common.llm_limit): every judge call routes through common.llm_call and takes a slot, so judge fan-out across concurrent compare calls is capped run-wide.
assignment="cyclic" (CyclicJudge, arXiv:2603.01865) gives each comparison exactly one judge, cycling through the panel so every judge covers an equal share of the run. Judge bias cancels in expectation over many comparisons at single-judge cost, and the assigned judge still runs both orderings when swap is on. Per-pair winners are single-judge opinions — roll them up with build_report and read run-level rates. The rotation runs over the deduplicated panel (duplicate judge entries add weight only under "all"), repetitions still applies to the single assigned judge, and the cursor lives on the comparator: a reused comparator continues where the previous run stopped, so exact balance holds per freshly built comparator, not per run. Shuffle your pairs first if their order is meaningful. Requires min_successful_judges=1.
Usage:
from evaluatorq import llm_jury_pairwise
comparator = llm_jury_pairwise(
criteria="The answer is accurate, complete, and directly addresses the question.",
judges=["anthropic/claude-sonnet-4-6", "openai/gpt-5.6-luna"],
)
comparison = await comparator.compare(
question="What is the capital of France?",
response_a="The capital of France is Paris.",
response_b="The capital of France is Berlin.",
)
print(comparison.winner)
AgentResponse ¶
Bases: BaseModel
Structured response from a target agent as an ordered list of output messages.
Each item in output is a TextOutputItem, ToolCallOutputItem, or ReasoningOutputItem, preserving the order in which they were produced. Item type discriminators align with the OpenResponses intermediate data format (output_text, function_call, reasoning).
Accessors
.text — all TextOutputItem contents concatenated, or "" if none .tool_calls — list of ToolCallOutputItem filtered from output in order .refusal — the first provider refusal, when the response contains one
tool_calls property ¶
Return the tool call items from .output in order.
from_output_items classmethod ¶
Build an AgentResponse from bare Responses output items, with no usage or metadata.
from_openresponses classmethod ¶
Build a full AgentResponse from a Responses API response object.
Single parse path for the OpenResponses wire format. Populates output + usage + model + finish_reason + response_id.
usage is None when the response carries no usage block — callers must distinguish "no usage reported" from "zero tokens used" so cost reports stay honest. usage.calls is intentionally left at 0: this is a pure parse; per-call-site accounting (calls=1 / +1) is applied by callers to the returned object's usage.
BTFit ¶
Bases: BaseModel
Result of a Bradley-Terry family fit.
skills class-attribute instance-attribute ¶
Latent skill per item, centred at mean 0
sigmas class-attribute instance-attribute ¶
sigmas: dict[str, float] = Field(
description='Per-judge discriminator; smaller = sharper/more reliable. Only RATIOS between judges in the same fit are meaningful (the log-normal prior anchors the scale loosely); do not compare absolute values across fits. Empty when fitted without judge_sigma.'
)
Per-judge discriminator; smaller = sharper/more reliable. Only RATIOS between judges in the same fit are meaningful (the log-normal prior anchors the scale loosely); do not compare absolute values across fits. Empty when fitted without judge_sigma.
ranking class-attribute instance-attribute ¶
Item IDs, best first
converged class-attribute instance-attribute ¶
converged: bool = Field(
description='True when the iterate stopped moving before the iteration cap'
)
True when the iterate stopped moving before the iteration cap
iterations class-attribute instance-attribute ¶
Optimizer iterations used
warnings class-attribute instance-attribute ¶
warnings: list[str] = Field(
default_factory=list, description='Degradations applied during the fit'
)
Degradations applied during the fit
reliability ¶
1 / sigma for a judge - the paper's unsupervised reliability signal.
BTSigmaAggregation ¶
Bases: BaseModel
Reliability-weighted rollup of a pairwise run (BT-sigma, arXiv:2602.16610).
Fitted unsupervised on the run's own reconciled votes: hard BT-sigma jointly infers the A-vs-B skill gap and a discriminator per judge, then re-derives each comparison's winner as a reliability-weighted vote instead of uniform plurality. Down-weights noisy judges; needs no labels. The weight depends on the path: on the pooled two-item fit it is 1/sigma; on a repetition run (repetition_consistency non-empty and at quorum) it is instead max(shrunk_consistency, 0.05), since within-datapoint consistency is the better reliability signal there.
Two caveats of the two-item collapse, both handled here:
- With only two items a judge's sigma is pinned by its own vote distribution, so a judge whose decisive votes are unanimous gets an arbitrarily small sigma that measures one-sidedness (e.g. position or verbosity bias), not reliability. Such judges are excluded from the weighting — they vote with a neutral weight, their sigma is dropped from
judge_sigmas, andfit_warningsnames them. p_a_beats_b/skill_gapcome from the pooled global fit, whilewinnerscome from the per-comparison weighted vote. These are two different estimators and need not agree; readp_a_beats_bas the run-level headline andwinnersas the per-row calls.
p_a_beats_b class-attribute instance-attribute ¶
p_a_beats_b: float = Field(
description='Fitted global probability that A beats B (logistic of the skill gap)'
)
Fitted global probability that A beats B (logistic of the skill gap)
skill_gap class-attribute instance-attribute ¶
Fitted skill difference s_A - s_B
judge_sigmas class-attribute instance-attribute ¶
judge_sigmas: dict[str, float] = Field(
default_factory=dict,
description='Per-judge discriminator; smaller = sharper/more reliable. Empty on single-judge fallback.',
)
Per-judge discriminator; smaller = sharper/more reliable. Empty on single-judge fallback.
repetition_consistency class-attribute instance-attribute ¶
repetition_consistency: dict[str, float] = Field(
default_factory=dict,
description='Per-judge SHRUNK within-datapoint repetition consistency in [0, 1] (RES-1251); when non-empty and at quorum, the winner weights came from THESE, not from the global-fit sigmas. Shrunk toward the panel mean, so it is a reliability weight, not raw self-agreement - a self-consistent judge reads below 1.0 unless the panel is. Measures self-agreement on repeated passes - NOT task difficulty and NOT overall judge quality.',
)
Per-judge SHRUNK within-datapoint repetition consistency in [0, 1] (RES-1251); when non-empty and at quorum, the winner weights came from THESE, not from the global-fit sigmas. Shrunk toward the panel mean, so it is a reliability weight, not raw self-agreement - a self-consistent judge reads below 1.0 unless the panel is. Measures self-agreement on repeated passes - NOT task difficulty and NOT overall judge quality.
repetition_consistency_raw class-attribute instance-attribute ¶
repetition_consistency_raw: dict[str, float] = Field(
default_factory=dict,
description='Per-judge RAW within-datapoint self-agreement in [0, 1] (RES-1251): neither shrunk nor failure-adjusted, published alongside repetition_consistency so the un-shrunk signal (1.0 = always agreed with itself on every completed pass) stays visible next to the weight derived from it.',
)
Per-judge RAW within-datapoint self-agreement in [0, 1] (RES-1251): neither shrunk nor failure-adjusted, published alongside repetition_consistency so the un-shrunk signal (1.0 = always agreed with itself on every completed pass) stays visible next to the weight derived from it.
winners class-attribute instance-attribute ¶
winners: list[str] = Field(
default_factory=list,
description='Reliability-weighted winner per comparison, aligned with the input order',
)
Reliability-weighted winner per comparison, aligned with the input order
a_win_rate class-attribute instance-attribute ¶
a_win_rate: float | None = Field(
description='Weighted A-wins over decisive comparisons; None if none decisive'
)
Weighted A-wins over decisive comparisons; None if none decisive
b_win_rate class-attribute instance-attribute ¶
b_win_rate: float | None = Field(
description='Weighted B-wins over decisive comparisons; None if none decisive'
)
Weighted B-wins over decisive comparisons; None if none decisive
tie_rate class-attribute instance-attribute ¶
Weighted-tie comparisons over all comparisons
inconclusive_rate class-attribute instance-attribute ¶
inconclusive_rate: float = Field(
description='Comparisons no judge decided, over all comparisons'
)
Comparisons no judge decided, over all comparisons
converged class-attribute instance-attribute ¶
converged: bool = Field(
default=True,
description='False when the fit stopped at the iteration cap; treat sigmas and weighted winners with suspicion then (the cap is also reported in fit_warnings)',
)
False when the fit stopped at the iteration cap; treat sigmas and weighted winners with suspicion then (the cap is also reported in fit_warnings)
fit_warnings class-attribute instance-attribute ¶
fit_warnings: list[str] = Field(
default_factory=list, description='Degradations applied during the BT fit'
)
Degradations applied during the BT fit
ClassifyOutcome ¶
Bases: BaseModel
ClassifyQuestion ¶
Bases: BaseModel
One question for the Orq router's /classify endpoint.
A classify model is not prompted: it is handed the material to judge (state) and a question about it. kind picks the answer shape — noul a probability, choice one of a labelled set, score a probability-weighted mean over ordered levels — and criteria describes the answer space for that shape. noul_threshold is read by run_judge, not sent: it turns the probability into the boolean verdict a jury counts.
ClassifyRequest ¶
Bases: BaseModel
Several questions about one state, answered in one /classify round trip.
DataPointComplete module-attribute ¶
DataPointDict ¶
Bases: _DataPointDictRequired
Dict representation of a DataPoint for type checking.
DataPointInput module-attribute ¶
Type alias for DataPoint that accepts both model instances and dicts.
DataPointResult ¶
Bases: BaseModel
DeploymentResponse dataclass ¶
Response from a deployment invocation.
EvaluationResultCellValue module-attribute ¶
EvaluatorParams ¶
Bases: BaseModel
Parameters for running an evaluation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data | The data to evaluate. A DatasetIdInput to fetch from Orq platform, an ExperimentInput to replay an experiment's recorded responses (requires inference=False), a TraceInput to import recorded trace conversations (requires inference=False), or a list of DataPoint instances/awaitables. | required | |
jobs | The jobs to run on the data. | required | |
evaluators | The evaluators to use. If not provided, only jobs will run. | required | |
datapoint_parallelism | Number of datapoints to process in parallel. Defaults to 10; set to 1 for sequential execution. Accepts the former name | required | |
llm_parallelism | Ceiling on in-flight LLM requests for the whole run, counted per request rather than per task. Defaults to 10; -1 disables it. Use this against a provider concurrency limit — one datapoint can issue many requests, so the datapoint count cannot be sized against one. | required | |
print_results | Whether to print results table to console. Defaults to True. Also accepts "print" as an alias. | required | |
description | Optional description for the evaluation run. | required | |
path | Optional path (e.g. "MyProject/MyFolder") to place the experiment in a specific project and folder on the Orq platform. | required | |
single_trace | Group every row of the run under one | required |
inference class-attribute instance-attribute ¶
When False, skip generation and evaluate the pre-recorded response in each row's messages column instead of running jobs.
None resolves from data: a replay source (ExperimentInput, TraceInput) is a recorded-output source by construction and resolves to False, anything else to True. Passing inference=True with a replay source still raises — the flag was never able to mean anything else there, and requiring it taught users the mode only by making the obvious call fail.
single_trace class-attribute instance-attribute ¶
When True, open one evaluatorq.run span for the whole run so every row shares a trace. Default False keeps each row's orq.job as its own root.
EvaluatorScore ¶
Bases: BaseModel
EvaluatorqResult module-attribute ¶
Type alias for evaluation results
ExperimentInput ¶
Bases: BaseModel
Input for sourcing pre-recorded responses from an Orq experiment.
Used with inference=False to re-run evaluators against the responses an earlier experiment already produced, without regenerating them.
experiment_id instance-attribute ¶
The experiment ID to load responses from. Read it off the experiment URL in the Orq UI (/experiments/<experiment_id>). The API refers to experiments as "spreadsheets", so you will also see this ID in /v2/spreadsheets/<id> routes.
run_id class-attribute instance-attribute ¶
A specific run ID (a "manifest" in the API). When omitted, the latest run is used. Every execution of an experiment creates a new run; open it from the experiment's run history to read its ID from the URL.
Job module-attribute ¶
Job function type - returns a JobReturn-shaped dict ('name', 'output', optional 'error')
JobResult ¶
Bases: BaseModel
JobReturn ¶
Bases: TypedDict
Job return structure.
error is optional and reports a failure the job handled rather than raised — None on success, the reason otherwise. A row whose error flattens to a non-empty string is counted in the summary table's Failed Jobs and fails check_pass_failures(treat_errors_as_failure=True). It keeps its output for diagnosis but its evaluators are skipped — scoring a transcript already known to be dead buys nothing and costs an LLM judge call per row. An omitted key and an explicit None are indistinguishable to the consumer — emitting None on success is a producer convention, so that a job that forgot the key cannot be mistaken for one that reported a clean run. A job that lets its failures raise omits the key. Note that Job types this return as dict[str, Any], so nothing type-checks a job against this shape.
JudgeStats ¶
Bases: BaseModel
Per-judge behaviour rolled up across a set of comparisons.
a_rate class-attribute instance-attribute ¶
a_rate: float | None = Field(
description="Share of the judge's decisive picks that went to A; None if it never picked a side"
)
Share of the judge's decisive picks that went to A; None if it never picked a side
b_rate class-attribute instance-attribute ¶
b_rate: float | None = Field(
description="Share of the judge's decisive picks that went to B; None if it never picked a side"
)
Share of the judge's decisive picks that went to B; None if it never picked a side
position_bias class-attribute instance-attribute ¶
position_bias: float = Field(
description='Flips over completed pairs (both orderings decisive); 0.0 when no pair was flippable'
)
Flips over completed pairs (both orderings decisive); 0.0 when no pair was flippable
tie_rate class-attribute instance-attribute ¶
Ties over total comparisons the judge saw
sigma class-attribute instance-attribute ¶
sigma: float | None = Field(
default=None,
description="BT-sigma discriminator (smaller = more reliable); set only when the report was built with aggregation='bt-sigma'",
)
BT-sigma discriminator (smaller = more reliable); set only when the report was built with aggregation='bt-sigma'
consistency class-attribute instance-attribute ¶
consistency: float | None = Field(
default=None,
description='SHRUNK within-datapoint repetition consistency in [0, 1] (RES-1251): judge reliability measured from repeated passes of the same prompt and shrunk toward the panel mean. On a bt-sigma run where repetition weighting was actually applied, this is the quantity the winner weights came from, surfaced next to sigma rather than leaving the reader to infer reliability from a pooled sigma that did not decide anything; bt_sigma.fit_warnings names the runs where the weighting was skipped, and there this number is published without having decided anything. It is a diagnostic of the recorded repeats and needs no fit, so it is filled under plurality too. Because it is shrunk, a self-consistent judge reads below 1.0 unless the whole panel is; see consistency_raw for the un-shrunk value. None when the run had no usable repeats for this judge.',
)
SHRUNK within-datapoint repetition consistency in [0, 1] (RES-1251): judge reliability measured from repeated passes of the same prompt and shrunk toward the panel mean. On a bt-sigma run where repetition weighting was actually applied, this is the quantity the winner weights came from, surfaced next to sigma rather than leaving the reader to infer reliability from a pooled sigma that did not decide anything; bt_sigma.fit_warnings names the runs where the weighting was skipped, and there this number is published without having decided anything. It is a diagnostic of the recorded repeats and needs no fit, so it is filled under plurality too. Because it is shrunk, a self-consistent judge reads below 1.0 unless the whole panel is; see consistency_raw for the un-shrunk value. None when the run had no usable repeats for this judge.
consistency_raw class-attribute instance-attribute ¶
consistency_raw: float | None = Field(
default=None,
description='RAW within-datapoint self-agreement in [0, 1] (RES-1251): 1.0 = this judge always agreed with itself on every completed pass. Neither shrunk nor failure-adjusted (both live in the reliability weight); published beside the shrunk consistency so the reader sees the un-shrunk signal, not only the weight derived from it. None when no usable repeats.',
)
RAW within-datapoint self-agreement in [0, 1] (RES-1251): 1.0 = this judge always agreed with itself on every completed pass. Neither shrunk nor failure-adjusted (both live in the reliability weight); published beside the shrunk consistency so the reader sees the un-shrunk signal, not only the weight derived from it. None when no usable repeats.
JudgedComparison ¶
Bases: BaseModel
One judge's preference on one item pair, in a fixed (a, b) frame.
p_a is the judge's preference probability that item_a beats item_b. Callers that only have categorical votes map them as A -> 1.0, B -> 0.0, tie -> 0.5. Position-bias symmetrisation (paper Eq. 4) is the caller's job: build p_a from both orderings, e.g. via the pairwise module's swap/reconcile machinery, before fitting.
judge class-attribute instance-attribute ¶
Judge model ID
item_a class-attribute instance-attribute ¶
First item ID in the canonical frame
item_b class-attribute instance-attribute ¶
Second item ID in the canonical frame
p_a class-attribute instance-attribute ¶
P(item_a beats item_b) according to this judge
weight class-attribute instance-attribute ¶
weight: float = Field(
default=1.0,
gt=0.0,
description='Multiplicity: how many identical judgements this record stands for. The fit cost is O(iterations x records), so collapsing repeats keeps large runs fast without changing the optimum.',
)
Multiplicity: how many identical judgements this record stands for. The fit cost is O(iterations x records), so collapsing repeats keeps large runs fast without changing the optimum.
JuryPreset ¶
Bases: BaseModel
A named panel: fixed judges, fixed aggregation, and reserves for retirement.
estimated_cost_per_1k class-attribute instance-attribute ¶
estimated_cost_per_1k: float = Field(
description='USD per 1,000 pointwise items at 1,500 input / 1,500 output tokens, uncached.'
)
USD per 1,000 pointwise items at 1,500 input / 1,500 output tokens, uncached.
seated_efforts ¶
The reasoning effort each judge runs at: its catalogue default.
A preset sends no effort, so each seat runs at the provider default. None means the catalogue names no effort parameter for that model.
llm_jury takes one reasoning_effort for the whole panel, so passing one overrides every seat's default; per-judge call settings are a schema change and its own ticket (RES-1347).
duplicated_lineages ¶
Lineages seated more than once, whose errors correlate.
Not an error. A panel may repeat a lineage deliberately, as Single-Provider Trio does with three OpenAI judges, but the diversity claim weakens and the reviewer should see it rather than count IDs by hand.
cost_per_1k ¶
The panel's $/1k recomputed from the captured billing rates.
estimated_cost_per_1k is the figure the docs publish; this is where it has to come from. A test asserts the two agree, so a repricing in the snapshot fails CI rather than quietly making a published table wrong.
MessageDict ¶
Bases: TypedDict
Chat message structure compatible with Orq SDK.
Output module-attribute ¶
Output type alias
PairwiseComparator ¶
A configured pairwise LLM jury. Call compare on an A/B pair.
compare async ¶
Run the panel over one A-vs-B comparison and return the reconciled verdict.
A side that carries a target-level error is never judged — comparing against a failed generation would score noise as a preference.
PairwiseComparison ¶
Bases: BaseModel
The panel's result for a single A-vs-B comparison.
winner class-attribute instance-attribute ¶
Consensus winner: 'A', 'B', 'tie', or 'inconclusive'
votes class-attribute instance-attribute ¶
Per-judge reconciled votes
token_usage class-attribute instance-attribute ¶
token_usage: TokenUsage | None = Field(
default=None,
description='Summed token usage across both orderings and any replacements',
)
Summed token usage across both orderings and any replacements
PairwiseReport ¶
Bases: BaseModel
Cross-comparison rollup of a pairwise run.
comparisons class-attribute instance-attribute ¶
Number of comparisons in the run
a_win_rate class-attribute instance-attribute ¶
a_win_rate: float | None = Field(
description='Consensus A-wins over comparisons decided A or B; None if none were decisive'
)
Consensus A-wins over comparisons decided A or B; None if none were decisive
b_win_rate class-attribute instance-attribute ¶
b_win_rate: float | None = Field(
description='Consensus B-wins over comparisons decided A or B; None if none were decisive'
)
Consensus B-wins over comparisons decided A or B; None if none were decisive
tie_rate class-attribute instance-attribute ¶
tie_rate: float = Field(
description='Comparisons whose consensus was a tie, over all comparisons'
)
Comparisons whose consensus was a tie, over all comparisons
inconclusive_rate class-attribute instance-attribute ¶
inconclusive_rate: float = Field(
description='Comparisons with no consensus (panel too degraded to decide), over all comparisons'
)
Comparisons with no consensus (panel too degraded to decide), over all comparisons
mean_agreement class-attribute instance-attribute ¶
mean_agreement: float | None = Field(
description='Mean inter-judge agreement (modal vote share) across comparisons; None if none were decisive'
)
Mean inter-judge agreement (modal vote share) across comparisons; None if none were decisive
per_judge class-attribute instance-attribute ¶
per_judge: list[JudgeStats] = Field(
default_factory=list, description='Per-judge behaviour breakdown'
)
Per-judge behaviour breakdown
bt_sigma class-attribute instance-attribute ¶
bt_sigma: BTSigmaAggregation | None = Field(
default=None,
description="Reliability-weighted rollup; set only when the report was built with aggregation='bt-sigma'",
)
Reliability-weighted rollup; set only when the report was built with aggregation='bt-sigma'
PairwiseVote ¶
Bases: BaseModel
One judge's reconciled verdict for a single A-vs-B comparison.
model class-attribute instance-attribute ¶
Judge model ID
vote class-attribute instance-attribute ¶
vote: Literal['A', 'B', 'tie'] | None = Field(
description="Reconciled vote: 'A', 'B', 'tie', or None if the judge abstained"
)
Reconciled vote: 'A', 'B', 'tie', or None if the judge abstained
flipped class-attribute instance-attribute ¶
flipped: bool = Field(
default=False,
description='True if the judge disagreed with itself across orderings',
)
True if the judge disagreed with itself across orderings
completed class-attribute instance-attribute ¶
completed: bool = Field(
default=True,
description='True if both orderings produced a decisive verdict, so a flip was actually possible',
)
True if both orderings produced a decisive verdict, so a flip was actually possible
replacement class-attribute instance-attribute ¶
replacement: bool = Field(
default=False,
description='True if this judge stood in for a failed configured judge',
)
True if this judge stood in for a failed configured judge
explanation class-attribute instance-attribute ¶
explanation: str = Field(
default='', description='Explanation from the reconciled decisive ordering'
)
Explanation from the reconciled decisive ordering
observations class-attribute instance-attribute ¶
observations: list[RepetitionObservation] = Field(
default_factory=list,
description='Canonicalized per-repetition votes across both orderings (RES-1251); empty on runs saved before repetition capture existed',
)
Canonicalized per-repetition votes across both orderings (RES-1251); empty on runs saved before repetition capture existed
repetition_failures class-attribute instance-attribute ¶
repetition_failures: int = Field(
default=0,
ge=0,
description='Repetition passes that raised an error OR returned an off-contract value, summed over both orderings (a clean abstention is not a failure)',
)
Repetition passes that raised an error OR returned an off-contract value, summed over both orderings (a clean abstention is not a failure)
RepetitionObservation ¶
Bases: BaseModel
One raw repetition pass of one judge in one ordering, canonicalized.
verdict is in the ORIGINAL A/B orientation regardless of ordering: a 'B' returned under the swapped ordering is recorded as 'A' here, so the observations are directly comparable across orderings. None is a pass that produced no usable verdict: a genuine abstention, an error, or an off-contract value. The per-vote repetition_failures count says how many of those Nones were errors or off-contract (as opposed to clean abstentions); those two lower reliability while an abstention does not (RES-1251). Nothing is silently collapsed.
ordering class-attribute instance-attribute ¶
ordering: Literal['ab', 'ba'] = Field(
description="'ab' = original slots, 'ba' = swapped ordering"
)
'ab' = original slots, 'ba' = swapped ordering
repetition class-attribute instance-attribute ¶
Call order within the ordering, 0-based
verdict class-attribute instance-attribute ¶
verdict: Literal['A', 'B', 'tie'] | None = Field(
default=None, description='Canonicalized verdict; None = abstained or failed pass'
)
Canonicalized verdict; None = abstained or failed pass
explanation class-attribute instance-attribute ¶
explanation: str | None = Field(
default=None, description="This pass's own reasoning; None when it produced no text"
)
This pass's own reasoning; None when it produced no text
raw_output class-attribute instance-attribute ¶
raw_output: dict[str, Any] | None = Field(
default=None,
description='Validated structured judge output in the ordering sent to the provider.',
)
Validated structured judge output in the ordering sent to the provider.
Scorer module-attribute ¶
ScorerParameter ¶
Bases: TypedDict
Parameters passed to a scorer function.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data | The data point being evaluated. | required | |
output | The output produced by the job for the data point. | required | |
row | Zero-based dataset index of the data point. Present when the scorer runs inside | required |
ThreadConfig ¶
Bases: TypedDict
Thread configuration for conversation tracking.
Trace ¶
Bases: BaseModel
A normalized Orq trace exchange and its source metadata.
input_messages and output_messages retain their source-side boundaries. messages is a derived transcript for consumers that need a single conversation, with only the largest exact overlap removed.
TraceInput ¶
Bases: BaseModel
Describe one exact trace/span, one trace, or a bounded trace search.
The three modes are exclusive and the validator enforces that, limit included: in trace mode it is never read, so accepting it would answer limit=50 with exactly one trace and no warning. Read the resolved count off query_limit rather than the field.
start_time / end_time are stored timezone-aware: a naive value is read as UTC rather than as the host's local time, so one query means one window wherever it runs.
limit class-attribute instance-attribute ¶
Query mode only. None means "not set", which is what lets the validator tell an explicit limit= apart from the default it would otherwise assume.
bt_sigma_aggregation ¶
Fit hard BT-sigma over a run's reconciled votes and re-derive winners.
The two "items" are the run's A and B sides; every decisive reconciled vote is one comparison between them. Reconciliation has already symmetrised position bias (both orderings per judge), satisfying the model's commutativity requirement. Categorical votes make hard BT-sigma the natural variant, which is also the paper's most robust one under inconsistency.
Identical votes are collapsed into weighted records before the fit (a judge has at most three distinct judgements here: A, B, tie), so the fit cost stays flat no matter how many comparisons the run holds.
Judges whose decisive votes are unanimous are excluded from the reliability weighting: in the two-item collapse their sigma measures one-sidedness, not reliability (see BTSigmaAggregation). They vote with the median weight of the remaining judges (or uniformly when no judge remains), and fit_warnings names them.
build_report ¶
build_report(
comparisons: Sequence[PairwiseComparison], *, aggregation: str = 'plurality'
) -> PairwiseReport
Roll a set of pairwise comparisons up into headline and per-judge metrics.
aggregation='plurality' (default) keeps the existing uniform plurality consensus. aggregation='bt-sigma' additionally fits hard BT-sigma over the run and attaches the reliability-weighted rollup (report.bt_sigma) plus each judge's discriminator (JudgeStats.sigma); the headline plurality rates are unchanged so the two aggregations stay comparable.
JudgeStats.consistency / consistency_raw are filled under BOTH aggregations: they are read straight off the recorded repetition observations and need no Bradley-Terry fit, so gating them on bt-sigma would report "no usable repeats" for a plurality run that has them.
exact_match_evaluator ¶
Creates an evaluator that checks if the output exactly matches the expected output. Uses the data.expected_output from the dataset to compare against.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
case_insensitive | bool | Whether the comparison should be case-insensitive | False |
name | str | Optional name for the evaluator | 'exact-match' |
Returns:
| Type | Description |
|---|---|
Evaluator | An Evaluator that checks if output exactly matches expected output |
fetch_traces async ¶
fetch_traces(
source: TraceInput,
*,
api_key: str | None = None,
base_url: str | None = None,
http_client: AsyncClient | None = None,
) -> list[Trace]
Fetch Orq traces and normalize each one to a canonical trace.
One unreadable trace becomes a Trace with import_error set; other traces in the batch remain available. No trace is silently discarded. Partition the result with partition_traces rather than filtering by hand — every surface is meant to report import failures the same way.
The HTTP calls retry through common.retry.with_retry — including the 429s and 5xx that raise_for_status raises as httpx.HTTPStatusError — and the shared httpx.AsyncClient adds no retry layer of its own. Raw httpx rather than the Orq SDK because traces.search requires a from_/to window that TraceInput does not, and its span models would need a second parser.
fit_bt ¶
fit_bt(
comparisons: Sequence[JudgedComparison],
*,
judge_sigma: bool = True,
hard: bool = False,
) -> BTFit
Fit the Bradley-Terry family by MLE on the given comparisons.
judge_sigma adds the per-judge discriminator (BT-sigma). hard binarizes preferences first (hard BT / hard BT-sigma) - the paper's most robust variant when judges are highly inconsistent.
Degradations (recorded in warnings, never raised):
- a single judge with
judge_sigma=Truefalls back to plain BT - a lone sigma is absorbed into the skill scale and carries no information (paper section 3.3); - a tiny ridge keeps disconnected graphs / perfect separation finite.
get_preset ¶
Look a preset up by name, listing the alternatives when there is no match.
llm_concurrency_limit ¶
Bound concurrent LLM requests within this block.
None sets nothing: enclosing limits apply, else the shared default of DEFAULT_LLM_PARALLELISM. -1 (UNBOUNDED) adds no ceiling; at the top level it disables the default. Nested limits stack, so an inner limit cannot raise or disable an outer ceiling.
Set once per run, before fanning out: tasks created inside inherit the limit, tasks created before it do not. Each entry has its own budget, while nested entries also share their ancestors' budgets.
Enterable with with or async with — the entry points wrap at whichever block they already open, and reindenting theirs to match would rewrite hundreds of untouched lines. A sync with may span several asyncio.run calls.
An enclosing limit is left alone when max_concurrent is None.
llm_slot async ¶
Hold one slot from every active finite LLM budget.
Wrap only the request itself — holding a slot across parsing or judging shrinks the budget without reducing load on the provider.
orq_evaluator ¶
orq_evaluator(
*,
evaluator_id: str,
model: str | None = None,
client: _OrqClient | None = None,
api_key: str | None = None,
base_url: str | None = None,
) -> Evaluator
Create an evaluatorq scorer backed by an Orq evaluator.
evaluator_id is required and keyword-only. It identifies an evaluator that already exists in Orq. The optional model is an invocation-time override for LLM-based evaluators; it is not needed for deterministic built-in evaluators, and a configured LLM evaluator can use its Orq model when this is omitted. When omitted, model is not sent to the SDK at all.
The evaluatorq result name is derived, never hand-named, from everything that varies which judge actually runs: orq:<evaluator_id> when model is omitted, orq:<evaluator_id>@<model> when it is set. This keeps two orq_evaluator calls that differ only by model — the standard way to compare judges — distinguishable in the results table and the uploaded experiment; a single fixed orq:<evaluator_id> name would have merged them into one indistinguishable stream.
client accepts an already-built Orq SDK client (or a test double shaped like one). When omitted, one is resolved from api_key or ORQ_API_KEY — and base_url or ORQ_BASE_URL, for a self-hosted deployment — the first time the scorer actually runs, and reused for every later datapoint scored by this same orq_evaluator(...) call. Resolving lazily — inside the scorer, not at factory time — preserves the property that constructing an evaluator that is never invoked does no work and raises nothing; caching that resolution after the first call keeps evaluate()'s default datapoint_parallelism (10 concurrent rows, or more across a large run) from building one Orq(...) client and one unclosed connection pool per row, and from deferring a missing ORQ_API_KEY or missing [orq] extra until after every row has already run and billed.
The scorer sends the latest user turn as query, evaluatorq's job output as output, expected_output as the optional human/reference answer, and preserves prior assistant/tool trajectory messages. Orq invocation trace IDs remain in EvaluationResult.raw_output for local use; evaluatorq's Experiment uploader intentionally strips raw evaluator output. The Orq SDK owns retries for this call; evaluatorq does not add another retry layer.
repetition_consistency ¶
Per-judge within-datapoint repetition consistency, in [0, 1] (RES-1251).
The unit of analysis is one (judge, comparison, ordering) group: repeated passes of the SAME prompt. A group needs >= 2 decisive repetitions; its consistency is the mean pairwise agreement among them, then discounted by the share of that vote's passes that errored or came back off-contract (repetition_failures) so a flaky judge scores below a clean one; a genuine abstention is not counted as a failure and does not penalise. A comparison contributes at most ONE observation per judge (its groups averaged), so repetition count never multiplies a judge's evidence, and different datapoints are never compared to each other - which is exactly why this is interpretable as judge reliability where the global two-item fit was not.
Empty dict when no judge has any qualifying group (e.g. repetitions=1 or a legacy run without observations).
This is the SHRUNK reliability weight, not raw self-agreement: at the recommended R=2 a group's agreement is only 0.0 or 1.0, so a judge with one lucky group would otherwise carry the same weight as a judge with many. Each judge's mean is shrunk toward the panel mean by its evidence count (empirical-Bayes, _SHRINKAGE_PSEUDO_OBS pseudo-observations), so thin evidence is pulled to neutral and cannot dominate a run. A perfectly self-consistent judge therefore reads below 1.0 unless the whole panel is; repetition_consistency_raw publishes the un-shrunk mean beside it (RES-1251).
repetition_consistency_raw ¶
Per-judge RAW within-datapoint self-agreement mean (RES-1251).
1.0 = a judge that always agreed with itself on repeated passes of the same prompt. This is PURE agreement, not failure-adjusted and not shrunk: a judge that agreed with itself on every completed pass reads 1.0 even if some passes failed (the failure discount and the shrinkage both live in the repetition_consistency weight, not here). Published beside that weight so the raw signal stays visible. Empty dict when no judge has a qualifying group.
run_classify async ¶
run_classify(
*,
client: AsyncOpenAI,
model: str,
cfg: LLMCallConfig,
request: ClassifyRequest,
span_attributes: dict[str, Any] | None = None,
) -> ClassifyOutcome
Call /classify once for all questions in request and validate the reply.
The caller owns retry policy. This helper owns the classify span, timeout, wire response validation, and mapping provider failures into ClassifyOutcome.
run_pairwise async ¶
run_pairwise(
*,
judge_fn: PairwiseJudgeFn,
panel: Sequence[str],
response_a: Any,
response_b: Any,
swap: bool = True,
repetitions: int = 1,
replacement_judges: Sequence[str] | None = None,
min_successful_judges: int = 1,
propagate_errors: bool = False,
) -> PairwiseComparison
Run a panel over one A-vs-B comparison and reconcile it into a verdict.
Each judge runs through the shared run_jury. When swap is on (default) every judge is also run with A and B exchanged; the two verdicts are un-swapped and passed through reconcile_pair, so a judge that just follows slot order abstains and is recorded as a flip.
Both orderings run concurrently. Replacement judges are promoted at the pair level — a stand-in is run in both orderings so it can cast a real reconciled vote, and a primary that fails in either ordering is what it stands in for. The winner is the plurality of the reconciled votes, or 'inconclusive' when fewer than min_successful_judges cast a decisive reconciled vote.
Concurrency is bounded by the run-scoped LLM ceiling (see common.llm_limit): every judge call routes through common.llm_call and takes a slot, so both orderings and the replacement pass share one run-wide cap rather than a per-comparison one.
Tracing: the whole comparison is ONE orq.pairwise_jury span (RES-985). Each ordering drives _run_jury_core rather than run_jury so it doesn't mint a second jury span whose aggregates describe half a comparison; every judge span hangs off this one, tagged judge.label_swapped.
PRESETS module-attribute ¶
PRESETS: MappingProxyType[str, JuryPreset] = MappingProxyType({
(preset.name): preset
for preset in (
BALANCED_TRIO,
STRONG_JURY,
OPEN_WEIGHT_PORTABLE,
EU_REGION,
SINGLE_PROVIDER_TRIO,
)
})