Agent simulation output reference¶
A SimulationResult is the record of one simulated conversation between a persona (an LLM playing a user with set traits) and the target (the agent under test), scored by a judge (an LLM that checks the conversation against the scenario). It holds the transcript, the judge's verdict, and per-turn scores. simulate() returns these as a list. A saved run wraps them in a SimulationRun, which adds the run's configuration, evaluator averages and the datapoints needed to replay it.
Read this page when you parse simulation output in code, in CI or from an agent. It does not cover writing personas and scenarios or reading the dashboard; for those, see Agent Simulation. For every field and its type, see the Python API reference.
A scenario is the situation and goal the persona brings. A datapoint is one persona paired with one scenario. A criterion is a behaviour the scenario says must happen or must not happen. An evaluator turns a finished conversation into a numeric score.
Where the output lands¶
| You ran | What you get | Model |
|---|---|---|
simulate() or generate_and_simulate() | the return value | list[SimulationResult] |
the same with save=True | .evaluatorq/sim-runs/<name>_<YYYYmmdd-HHMMSS>.json, or report_path= when set | SimulationRun |
evaluatorq sim simulate or sim run (saves by default; --no-save skips) | .evaluatorq/sim-runs/<name>_<YYYYmmdd-HHMMSS>.json | SimulationRun |
the CLI with --report out.json | out.json | SimulationRun |
the CLI with --results out.jsonl | one result per line | SimulationResult |
The SDK does not save by default and the CLI does. The .evaluatorq/sim-runs/ files are what evaluatorq sim runs, the dashboard and previous_run= read.
The shape¶
SimulationRun:
run_name, created_at, mode, run_id, experiment_url
target_kind, target, target_model, agent_info, orq_base_url, max_turns
evaluator_names, total_results
scorer_averages: dict[str, float] # mean of each evaluator over the run
token_usage_total # whole run, including generation and summary
executive_summary, recommendations, applied_suggestions
datapoints: list[SimulationDatapoint] # the exact cases, for replay
replay_version
results: list[SimulationResult] # one per conversation
messages # the transcript
terminated_by, reason, turn_count
goal_achieved, goal_completion_score
criteria_results, criteria_verified, rules_broken
turn_metrics: list[TurnMetrics] # per-turn judge scores and tokens
token_usage, token_usage_known
evaluator_details # structured evaluator output
metadata # persona, scenario, criteria, evaluator scores
thread_id, response_traces
Reading one result¶
This result is trimmed from a saved run. The transcript, turn metrics and token counts are cut. The run predates audited and thread_id, so those two are filled in by hand to show what a current run writes.
{
"terminated_by": "max_turns",
"reason": "Maximum turns (4) reached",
"goal_achieved": false,
"goal_completion_score": 0.0,
"rules_broken": [],
"turn_count": 4,
"criteria_results": {"Agent greets the user politely and responds to the greeting": true, "Agent asks for unnecessary personal information": true},
"criteria_verified": true,
"metadata": {
"persona": "Curious Office Receptionist",
"scenario": "Friendly Greeting and Weather Check",
"datapoint_id": "dp_3e44c06b57af",
"criteria_meta": [
{"id": "criteria_0", "description": "Agent greets the user politely and responds to the greeting", "type": "must_happen", "passed": true, "audited": true},
{"id": "criteria_2", "description": "Agent asks for unnecessary personal information", "type": "must_not_happen", "passed": true, "audited": true}
],
"evaluator_scores": {"goal_achieved": 0.0, "criteria_met": 1.0}
},
"thread_id": "<run_id>:0"
}
terminated_by says why the conversation stopped:
| Value | Means |
|---|---|
judge | The judge ended it: the goal was reached, a must-not-happen criterion occurred, or the judge's own reply could not be parsed and it stopped for safety. reason says which. |
max_turns | It hit the turn cap without the judge ending it. The goal was not reached in the turns allowed, so goal_achieved is usually false; whether that fails your gate is your call. Criteria are still scored. |
timeout | A target call, or the whole conversation, ran out of time. reason describes the failure; metadata["error_type"] identifies its exception type. |
error | A target call failed after its retries, or the run itself raised (the simulated user, the judge, or setup before the first turn). reason describes the failure; metadata["error_type"] identifies its exception type. |
The criteria verdicts come in three places. metadata["criteria_meta"] is the one to read: one entry per criterion, keyed by a stable id (criteria_0, criteria_1, and so on), with its text, type, pass or fail, and the judge's evidence. criteria_results maps each criterion's description to pass or fail, so two criteria with the same description collapse into one key. rules_broken lists the ids of criteria that failed when criteria_verified is true; otherwise it is the judge's own free-text list, passed through.
passed: true means the criterion was met, whatever its type. For a must-happen criterion, the behaviour occurred. For a must-not-happen criterion, the behaviour did not occur. In the example, "Agent asks for unnecessary personal information": true means the agent did not ask.
Check criteria_verified before trusting criteria_results. When it is false, the judge never returned a per-criterion audit, and every verdict fell back to the free-text rules_broken list. That fallback cannot fail a must-happen criterion, so an agent that never did the required thing still shows it as passed. Treat those verdicts as unknown. The same applies to one criterion whose criteria_meta entry has "audited": false and "passed": true. audited is missing on runs saved before the field existed; that means "not recorded", not "unaudited".
metadata is a plain dict, so its keys are a convention rather than a schema. The keys evaluatorq writes:
| Key | Holds |
|---|---|
persona, scenario | the names of the persona and scenario |
persona_traits | the persona's patience, assertiveness, politeness, technical level, style and background |
scenario_goal, scenario_context | the scenario's goal and context |
criteria_meta | one entry per criterion: id, description, type, passed, audited, evidence |
datapoint_id | the datapoint this conversation ran |
evaluator_scores | evaluator name to numeric score, for this conversation |
evaluator_errors | evaluator name to the reason its score was unusable, when one failed |
target_model | the model the target used, when the client knows it |
error, error_type | why the conversation failed, on error and timeout results |
timeout | the time limit in seconds, when the conversation hit it |
token_usage_unknown | true when a failed conversation's usage could not be collected |
turn_metrics has one entry per turn with the judge's response_quality, hallucination_risk, tone_appropriateness and factual_accuracy (each 0 to 1, or null when not scored), its reasoning, and that turn's token usage.
token_usage_known is false when usage could only be partly collected. Treat token_usage as unknown then, not as a cheap run.
Reading the run¶
scorer_averages is the mean of each evaluator's metadata["evaluator_scores"] over the results that have one. A conversation where an evaluator failed is left out of that evaluator's average rather than counted as zero, so a run with many failures can still show a high average. Count metadata["evaluator_errors"] to see how many were left out.
datapoints holds the persona and scenario objects each conversation ran, which results keeps only by name. It is null on runs saved before replay existed, and those runs cannot be replayed.
recommendations lists suggested fixes for failed conversations, or is null when none were generated. Each entry points back to its result with result_index. applied_suggestions holds the suggestion strings already applied to the agent, so the dashboard marks them and a later apply skips them.
agent_info is a snapshot of an Orq agent's configuration at run time: key, id, role, description, model and tools, never its instructions. It is null for targets that are not Orq agents.
Load a saved run¶
This prints each conversation in the newest saved run, and flags unverified criteria:
from pathlib import Path
from evaluatorq.simulation import SimulationRun
latest = max(Path('.evaluatorq/sim-runs').glob('*.json'), key=lambda p: p.stat().st_mtime)
run = SimulationRun.model_validate_json(latest.read_text())
print(f'{run.run_name}: {run.total_results} conversations, averages {run.scorer_averages}')
for result in run.results:
criteria = 'unverified' if result.criteria_verified is False else result.criteria_results
print(f'{result.metadata.get("persona")} / {result.metadata.get("scenario")}: goal={result.goal_achieved} stop={result.terminated_by.value} criteria={criteria}')
To load a --results out.jsonl file from evaluatorq sim simulate instead, validate each line as a SimulationResult:
from pathlib import Path
from evaluatorq.simulation import SimulationResult
results = [SimulationResult.model_validate_json(line) for line in Path('out.jsonl').read_text().splitlines() if line]
Without Python, jq reads a saved run:
latest=$(ls -t .evaluatorq/sim-runs/*.json | head -1)
jq '{run_name, total_results, scorer_averages}' "$latest"
jq -r '.results[] | select(.goal_achieved == false) | "\(.metadata.persona) / \(.metadata.scenario): \(.reason)"' "$latest"
Gate CI on a run¶
This exits non-zero when any conversation errored or timed out, missed its goal, failed a criterion, or has a criterion the judge never audited:
import sys
from pathlib import Path
from evaluatorq.simulation import SimulationRun
latest = max(Path('.evaluatorq/sim-runs').glob('*.json'), key=lambda p: p.stat().st_mtime)
run = SimulationRun.model_validate_json(latest.read_text())
problems = []
for i, r in enumerate(run.results):
name = f'#{i} {r.metadata.get("persona")} / {r.metadata.get("scenario")}'
if r.terminated_by.value in ('error', 'timeout'):
problems.append(f'{name}: {r.terminated_by.value}: {r.reason}')
elif not r.goal_achieved:
problems.append(f'{name}: goal not achieved ({r.terminated_by.value}: {r.reason})')
for c in r.metadata.get('criteria_meta') or []:
if not c['passed']:
problems.append(f'{name}: failed {c["id"]} {c["description"]}')
elif c.get('audited') is False:
problems.append(f'{name}: {c["id"]} never audited')
if r.criteria_verified is not True:
problems.append(f'{name}: criteria not verified')
print('\n'.join(problems) or 'ok')
sys.exit(1 if problems else 0)
criteria_meta is the complete list of failed criteria, so rules_broken adds nothing to this check. criteria_verified is not True also fails runs saved before the field existed, where it is null: a gate that must never pass silently should treat "not recorded" as unverified. Drop the goal_achieved branch if your scenarios are exploratory and a conversation that runs out of turns is acceptable; keep the terminated_by branch, or a dead target passes.
Runs from older versions¶
Fields added after a run was saved load as their default: criteria_verified, thread_id, response_traces, run_id, datapoints, token_usage_total and orq_base_url are null or empty. A null there means "not recorded". criteria_verified: null in particular does not mean verified.