Skip to content

Evaluation Reference

Everything evaluatorq() accepts, and the patterns built on top of it. If you have not run an evaluation yet, start with Getting Started.

evaluatorq()

async def evaluatorq(
    name: str,
    params: EvaluatorParams | dict[str, Any] | None = None,
    *,
    data: DatasetIdInput | ExperimentInput | TraceInput | Sequence[Awaitable[DataPoint] | DataPointInput] | None = None,
    jobs: list[Job] | None = None,
    evaluators: list[Evaluator] | None = None,
    datapoint_parallelism: int = 10,
    llm_parallelism: int | None = None,
    print_results: bool = True,
    description: str | None = None,
    path: str | None = None,
    inference: bool = True,
) -> EvaluatorqResult
Parameter Type Default Description
data list[DataPoint \| dict] | list[Awaitable[DataPoint]] | DatasetIdInput | ExperimentInput | TraceInput required Data to evaluate — local rows (a DataPoint or a plain dict with the same keys), an Orq dataset, an existing experiment, or recorded trace output
jobs list[Job] | None required when inference=True Jobs to run on each data point; omitted and ignored when inference=False
evaluators list[Evaluator] | None None Evaluators that score job outputs
datapoint_parallelism int (≥1) 10 Number of concurrent datapoints. The former name parallelism still works, deprecated
llm_parallelism int (≥1, or -1) | None None Ceiling on in-flight LLM requests for the whole run. None keeps an enclosing limit or applies the default of 10; -1 adds no cap and disables the default only at the top level
print_results bool True Display the progress and results table
description str | None None Optional evaluation description
path str | None None Path for organizing results on the Orq dashboard (e.g. "Project/Category")
inference bool True Run the jobs; set False to score data that already has outputs

Parameters can also be passed positionally as an EvaluatorParams model or a plain dict — the three forms below are equivalent:

await evaluatorq("my-eval", data=[...], jobs=[...], datapoint_parallelism=5)
await evaluatorq("my-eval", {"data": [...], "jobs": [...], "datapoint_parallelism": 5})
await evaluatorq("my-eval", EvaluatorParams(data=[...], jobs=[...], datapoint_parallelism=5))

Full type signatures live in the API Reference.

Jobs

The @job() decorator

@job() names a job. The name shows up in the results table, in traces, and — crucially — in error messages:

from evaluatorq import job

@job("risky-job")
async def risky_operation(data: DataPoint, row: int):
    return await potentially_failing_operation(data)

# Error output: "Job 'risky-job' failed: <error details>"
# Without @job:  "<error details>"

It also wraps plain callables, which is handy for one-liners:

uppercase_job = job("uppercase", lambda data, row: data.inputs["text"].upper())
word_count_job = job("word-count", lambda data, row: len(data.inputs["text"].split()))

Multiple jobs per data point

Every job runs against every data point, which is how you compare variants (two prompts, two models, preprocessing on and off) on identical inputs:

await evaluatorq(
    "multi-job-eval",
    data=[...],
    jobs=[preprocessor, analyzer, transformer],
    evaluators=[...],
)

Reporting a failure the job handled

A job that lets a failure raise needs nothing: process_job records it and the row counts as failed. A job that catches its failure to keep the rest of the batch alive — a target that answered 401, a simulation the runner ended in error — must say so, because a returned dict looks like a clean run:

async def resilient_job(data: DataPoint, row: int) -> dict:
    try:
        answer = await call_my_agent(data.inputs["text"])
    except MyAgentError as exc:
        # Keeps the partial output for diagnosis, and still fails the row.
        return {"name": "my-agent", "output": None, "error": str(exc)}
    return {"name": "my-agent", "output": answer, "error": None}

Emit error on every path, None on success: an omitted key and a clean run are indistinguishable, so a job that forgets the key on one branch reports a dead target as a passing one. A row with a non-empty error is counted in the summary table's Failed Jobs and keeps its output for diagnosis, but its evaluators are skipped — scoring a transcript you already know is dead buys nothing and costs an LLM judge call per row. check_pass_failures(results, treat_errors_as_failure=True) is what turns it into a CI failure; the default (False) gates on evaluator pass_ alone.

This is the raw-dict job contract. @job() wraps a function's return value into {"name", "output"}, so an error key returned from a decorated function lands inside output, where a judge reads it as the target's failure rather than the runner reading it as the row's. A decorated job reports a row failure by raising.

Calling an Orq deployment from a job

When the thing you are evaluating is an Orq deployment, call it from inside the job. evaluatorq.deployment wraps the Orq SDK in two async functions: deployment() returns a DeploymentResponse carrying both content (the extracted text) and raw (the untouched SDK response), and invoke() is the same call returning just the text. One Orq client is created lazily on first use and reused for the process.

This needs ORQ_API_KEY. A missing key raises ValueError at the first call, naming the variable.

from evaluatorq import DataPoint, job
from evaluatorq.deployment import invoke


@job("orq-deployment-job")
async def my_job(data: DataPoint, _row: int) -> str:
    return await invoke("my-deployment", inputs=data.inputs)

inputs fills the deployment's template variables, so a summarizer whose prompt references {{text}} takes inputs={"text": "Long article..."}.

For a chat-style deployment, pass messages — and reach for deployment() when you want the raw response as well:

from evaluatorq.deployment import deployment

response = await deployment("chatbot", messages=[{"role": "user", "content": "Hello!"}])
print(response.content)  # extracted text
print(response.raw)      # full SDK response object

thread={"id": "conversation-123"} groups several calls into one conversation on the platform. context (routing attributes) and metadata are also accepted — see deployment() in the API reference for the full signature.

To red-team or simulate that same deployment rather than evaluate it, name it as a target instead: Targets › Orq-hosted.

Data sources

data accepts inline DataPoints, an Orq dataset, or awaitables that resolve to DataPoints — the last of which lets you stream rows in from a slow source without blocking the run:

async def get_data_point(i: int) -> DataPoint:
    await asyncio.sleep(0.01)  # e.g. a network fetch
    return DataPoint(inputs={"value": i})

await evaluatorq(
    "async-eval",
    data=[get_data_point(i) for i in range(1000)],
    jobs=[...],
)

For Orq-hosted datasets, pass data=DatasetIdInput(dataset_id="..."). That path requires ORQ_API_KEY — see Configuration. To score responses a past Orq experiment already produced, rather than generating new ones, pass ExperimentInput — see Replaying a past experiment.

Replaying a past experiment

Sometimes you want to score responses that an Orq experiment already produced instead of generating fresh ones — to try new evaluators against a past run, or to re-grade without paying for another round of generation. That is what no-inference mode does: pass inference=False and evaluators run against the recorded response in each row rather than calling any job.

The response source is chosen by the data argument to evaluatorq():

data value What it loads
list[DataPoint] In-memory datapoints.
DatasetIdInput(dataset_id=...) Rows from an Orq dataset (you supply or generate the responses).
ExperimentInput(experiment_id=..., run_id=...) The recorded responses from a past experiment run. Requires inference=False.

ExperimentInput sits alongside DatasetIdInput in the data union — it is not a dataset, it is a completed experiment run whose outputs get replayed.

Finding the IDs

Both IDs are read off the Orq UI:

  • experiment_id — the ID in the experiment URL, /experiments/<experiment_id>. The REST API calls experiments "spreadsheets", so the same ID appears in /v2/spreadsheets/<id> routes.
  • run_id — optional. Every execution of an experiment creates a new run (a "manifest" in the API). Open a run from the experiment's run history to read its ID from the URL. Omit it to replay the latest run.

Example

from evaluatorq import ExperimentInput, evaluatorq


async def run():
    await evaluatorq(
        "replay-past-experiment",
        data=ExperimentInput(experiment_id="<experiment_id>"),  # latest run
        evaluators=[my_evaluator],
        inference=False,
    )

Pin a specific run with run_id:

data=ExperimentInput(experiment_id="<experiment_id>", run_id="<run_id>")

ORQ_API_KEY must be set — the recorded rows are fetched from the Orq API. When inference=False, jobs is optional and ignored. Any row whose recorded response is missing or blank fails loudly rather than being silently skipped.

Evaluate recorded trace output

TraceInput is the async request for recorded trace data, and it has two mutually exclusive modes. Query mode selects a bounded batch by limit, time bounds (start_time/end_time), search text, or filters — the mode you reach for when you do not already have a trace ID. Trace mode takes one trace_id, with an optional exact span_id; span_id cannot be used without its trace, and limit cannot be combined with trace_id because trace mode never reads it. The shared importer returns normalized Trace objects and Trace.to_datapoint() supplies the evaluatorq row shape.

Query mode is the common path — score whatever ran recently without looking up an ID first:

from datetime import datetime, timedelta, timezone

from evaluatorq import TraceInput, evaluatorq, orq_evaluator

results = await evaluatorq(
    name='production-quality',
    data=TraceInput(
        limit=50,
        start_time=datetime.now(timezone.utc) - timedelta(days=1),
        search='checkout',
    ),
    evaluators=[orq_evaluator(evaluator_id='eval-1')],
)

filters is a list[dict[str, Any]] passed straight through to the platform's advanced trace filters — the same shape the Orq Traces UI produces. Each entry takes a field (the attr. prefix over the response's attributes path), an op, and values as a list of strings, even for a single value:

data=TraceInput(limit=50, filters=[{"field": "attr.orq.billing.cache_read_cost", "op": "gt", "values": ["0"]}])

Once you have a specific trace — from the Orq UI's trace URL, or orq traces search --from 24h --to now --query checkout -o json — trace mode replaces the whole query:

results = await evaluatorq(
    name='production-quality',
    data=TraceInput(trace_id='trace_123', span_id='span_456'),
    evaluators=[orq_evaluator(evaluator_id='eval-1')],
)

Evaluatorq skips jobs and sends the trace's recorded output to your evaluators — TraceInput resolves inference=False on its own, so there is no need to pass it. An exact span is parsed on its own; a trace or query selection starts from the latest eligible non-evaluator span and follows parents until it finds messages. Chat Completions, Responses, and OpenTelemetry GenAI messages are accepted from top-level span input/output, flat or nested span attributes, and span event attributes. A query that matches no traces logs a warning, and evaluatorq() raises rather than running a green zero-row evaluation.

orq_evaluator needs the Orq SDK, which the evaluatorq[orq] extra installs (uv add "evaluatorq[orq]") — a base evaluatorq install already carries it, but name the extra explicitly rather than relying on that. evaluator_id is required and keyword-only. It identifies an evaluator that already exists in Orq and determines the evaluatorq result name (orq:eval-1 when model is omitted, orq:eval-1@<model> when you set one); there is no separate name= override. model= is optional and overrides the configured model only for model-backed evaluators; deterministic built-ins do not need it. A missing trace response is an error, not a clean score, so check the returned results before trusting a rate.

Built-in evaluators

from evaluatorq import exact_match_evaluator, string_contains_evaluator

string_contains_evaluator()                        # case-insensitive by default
string_contains_evaluator(case_insensitive=False)  # case-sensitive
string_contains_evaluator(name="my-contains-check")  # custom name in the table
exact_match_evaluator()                            # case-sensitive by default

Both compare the job output against the data point's expected_output. For LLM-graded evaluators see LLM as a Jury; for structured, multi-dimensional scores see Structured Results.

Custom evaluators

An evaluator is a {"name": ..., "scorer": ...} pair whose scorer receives the data point and the job output and returns a score:

async def accuracy_scorer(params):
    data, output = params["data"], params["output"]
    score = calculate_score(output, data.expected_output)
    return {"value": score, "explanation": "High accuracy match" if score > 0.8 else "Partial match"}


await evaluatorq(
    "dataset-evaluation",
    data=DatasetIdInput(dataset_id="your-dataset-id"),
    jobs=[processor],
    evaluators=[{"name": "accuracy", "scorer": accuracy_scorer}],
)

Pass/fail and CI

An evaluator that returns pass_ turns the run into a gate:

async def quality_scorer(params):
    score = calculate_quality(params["output"])
    return {
        "value": score,
        "pass_": score >= 0.8,
        "explanation": f"Quality score: {score}",
    }

When any evaluator returns pass_: False, evaluatorq() returns the results; the library never exits the process. To make a script a CI gate, inspect the results and exit explicitly:

from evaluatorq.evaluatorq import check_pass_failures

results = await evaluatorq(...)
if check_pass_failures(results, treat_errors_as_failure=True):
    raise SystemExit(1)

treat_errors_as_failure=True also gates on rows that errored — a job that raised, a job that reported its own failure, and an evaluator whose every call failed. It defaults to False, which gates on evaluator pass_ alone, so a run whose target was dead throughout can pass a gate that leaves it off.

The results table gains a pass rate row — Pass Rate | 75% (3/4).

Controlling the run

Parallelism

await evaluatorq("parallel-eval", data=[...], jobs=[...], datapoint_parallelism=10)

datapoint_parallelism counts tasks, and the bounds nest: at most datapoint_parallelism datapoints run at once, and within each one a separate budget of the same size covers its jobs and then its evaluators. Ten datapoints each running ten evaluators is a hundred concurrent tasks, not ten.

Bounding LLM requests

Against a provider concurrency limit, size the request ceiling instead:

await evaluatorq("bounded-eval", data=[...], jobs=[...], llm_parallelism=20)

Unset, the ceiling is 10 concurrent requests, and that default also covers calls made outside evaluatorq(), such as a standalone run_pairwise(). Two unconfigured runs in the same event loop share that default cap of 10, protecting the provider from their combined load; a run with an explicit llm_parallelism= has its own budget. Pass -1 at the top level to disable the default ceiling. To bound code that is not an entry point, wrap it in async with evaluatorq.llm_concurrency_limit(n):. Nested limits stack: a block set to 5 inside a run set to 10 has a ceiling of 5, while a block set to 20 or -1 inside that run remains under the run's ceiling of 10.

This counts requests, not tasks, so it holds however the fan-out nests. It is a concurrency bound rather than a rate limit — ten slots against 10s calls is about 60 requests/minute, but the same ten slots become 300/minute if the provider speeds up to 2s.

Requests evaluatorq issues itself (judges, juries, simulation agents, the red-team pipeline) take a slot automatically. A job that calls a provider SDK directly is invisible to the budget unless you wrap it:

from evaluatorq import llm_slot

async def my_job(data_point, row_index):
    async with llm_slot():
        response = await client.chat.completions.create(...)
    return {"name": "my-job", "output": response.choices[0].message.content}

Wrap only the request — holding a slot across parsing shrinks the budget without reducing load on the provider.

red_team(), simulate(), generate_and_simulate() and generate() take the same argument, with the same meaning.

Organizing results on Orq

await evaluatorq(
    "my-evaluation",
    data=[...],
    jobs=[...],
    path="MyProject/Evaluations/Unit Tests",
)

path groups runs in the Orq dashboard — e.g. "Team/Sprint-42/Feature-X".

Documenting a run

await evaluatorq(
    "model-comparison",
    description="Compare GPT-4o vs Claude on customer support responses",
    data=[...],
    jobs=[...],
)

Suppressing terminal output

results = await evaluatorq("silent-eval", data=[...], jobs=[...], print_results=False)

for result in results:
    for job_result in result.job_results or []:
        print(f"{job_result.job_name}: {job_result.output}")

Where to next