Simulation (eq sim)¶
Agent simulation subcommand group. Always registered; the simulation extra (evaluatorq[simulation]) supplies the optional dependencies some of its commands need at runtime. sim is shorthand for convenience — the feature is agent simulation.
Three main verbs: generate (datapoints only), simulate (run against pre-built datapoints), run (generate then simulate in one shot).
The UI — eq dashboard
The way to browse saved simulation runs is the multi-run FastHTML dashboard, eq dashboard .evaluatorq/sim-runs (scopes to simulation) or eq dashboard (both stores). Passing a single JSON report file is an optional direct deep-link.
Going OpenAI-direct? Override the model roles
The default fast and smart models are openai/gpt-6-luna and openai/gpt-6-sol — provider-prefixed, because the default route is the Orq router, which resolves provider/model. With only OPENAI_API_KEY set, calls go straight to OpenAI, which does not know those ids and rejects them. Set EVALUATORQ_FAST_MODEL=gpt-6-luna and EVALUATORQ_SMART_MODEL=gpt-6-sol once, or pass the bare id on every model flag you use. The Models section of the Configuration page covers the roles.
eq sim run¶
Generate personas and scenarios, then run simulations.
Targets — provide exactly one:
| Flag | Description |
|---|---|
--target | agent:<key> or deployment:<key>. Bare values default to agent:<key>. |
--memory-entity | Memory entity_id sent with every agent:<key> (or bare <key>) target call, for agents with a memory store attached. Omit to mint a fresh id per conversation (parallel conversations never share memory); pass one to reuse a specific (e.g. seeded) entity, shared across the run. |
--vercel-url | Vercel AI SDK HTTP endpoint URL. |
--openai-model | OpenAI-compatible model name. Provider resolved from env: ORQ_API_KEY → Orq AI Router; OPENAI_API_KEY → OpenAI-compatible. |
| Flag | Type / Default | Description |
|---|---|---|
--agent-description | str \| None / None | Free-text description of the agent. May be omitted when --target is an Orq agent (fetched automatically). |
--name / -n | str \| None / unset | Run name. Unset unless passed: the SDK then saves the run as sim and names an uploaded experiment simulation-<timestamp>-<id>. |
--sim-model | str \| None / unset | Model for the user-simulator, the judge, persona/scenario/first-message generation, the recommendations pass and the executive summary. Unset, the simulated user, recommendations, executive summary and persona/scenario/first-message generation run on the fast role (openai/gpt-6-luna by default) and the judge runs on the smart role (openai/gpt-6-sol by default); a value here applies to all of them. Sets llm_config.model. |
--max-turns | int / 10 | Maximum conversation turns. Unset unless passed; the SDK default is 10. |
--datapoint-parallelism | int / 10 | Concurrent simulations. Unset unless passed; the SDK default is 10. --parallelism is a deprecated alias. |
--llm-parallelism | int / 10 | Ceiling on in-flight LLM requests for the whole run. -1 disables it. |
--num-personas | int / 5 | Number of personas to generate. |
--num-scenarios | int / 5 | Number of scenarios to generate. |
--persona-seed | str (repeatable) / None | Archetype seed for a persona, e.g. "angry retiree" (repeatable). Each seed becomes one persona the LLM fleshes out — overrides --num-personas. Omit to auto-generate. |
--scenario-seed | str (repeatable) / None | Situation seed for a scenario, e.g. "disputes a refund denial" (repeatable). Each seed becomes one scenario — overrides --num-scenarios. Omit to auto-generate. |
--generation-instructions | str / '' | Free-text steer applied to every generated persona AND scenario, e.g. "enterprise B2B buyers, replying in German". Stacks on top of any seeds and --edge-case-percentage. Empty leaves the built-in prompts unchanged. |
--target-reasoning-effort | str \| None / None | Reasoning effort pinned on the target agent under test (agent:<key> targets only). Distinct from the user-simulator's and judge's own reasoning effort, which comes from EVALUATORQ_REASONING_EFFORT — see Tuning. |
--evaluator | str (repeatable) / API defaults | Evaluator name(s). Repeatable. |
--save / --no-save | bool / True | Write the run to .evaluatorq/sim-runs/, with a run manifest. The SDK saves it after the recommendations and executive summary are attached. --no-save skips it. |
--recommendations / --no-recommendations | bool / True | Generate LLM remediation suggestions for failures, tied to their concrete cause. On by default; --no-recommendations skips the extra LLM call. Uses --sim-model. |
--datapoints / -d | Path \| None / None | Write generated datapoints to JSONL for reproducible re-runs. |
--results / -r | Path \| None / None | Path to write results JSONL, one SimulationResult per line. |
--report | Path \| None / None | Path to write full SimulationRun report JSON. The output reference describes this file and the --results rows. |
--report-md | Path \| None / None | Directory for an auto-named Markdown report. |
--report-html | Path \| None / None | Directory for an auto-named HTML report. |
--executive-summary / --no-executive-summary | bool / True | Generate an LLM narrative executive summary in the report. The SDK generates it before it saves the run. |
--yes / -y | bool / False | Skip interactive confirmation prompt. |
--verbose / -v | count / 0 | Increase verbosity. -v info; -vv debug. |
--quiet / -q | bool / False | Suppress non-error output. |
--config | PATH \| - / None | JSON file of generate_and_simulate() keyword arguments, or - to read it from stdin. See Driving the CLI from a config file; eq sim schema --command run prints the accepted shape. |
--llm-config | JSON / None | LLMCallConfig for every simulation-side LLM call as a JSON object. Merged field by field into "llm_config" from --config. --sim-model wins over its "model" when passed. |
--json | bool / False | Print the final SimulationRun as JSON on stdout. Progress and messages go to stderr; exit codes are unchanged. |
eq sim simulate¶
Run simulations from a pre-built datapoints JSONL file.
Targets — same three flags as eq sim run. Provide exactly one input source: --input (-i), --dataset-id, --experiment-id (optionally narrowed by --experiment-run-id), --from-run, or inline datapoints (or personas with scenarios) in --config.
There is no --target-reasoning-effort flag here — it is a eq sim run flag only. To pin the target's reasoning effort on a pre-built datapoint set, set "target_reasoning_effort" in --config, or call simulate() with target_reasoning_effort= from Python (see Tuning).
| Flag | Type / Default | Description |
|---|---|---|
--input / -i | Path \| None | Path to datapoints JSONL file. Mutually exclusive with the other input sources. |
--dataset-id | str \| None | Fetch datapoints from an Orq dataset instead of a local file. Requires ORQ_API_KEY. |
--experiment-id | str \| None | Fetch datapoints from an Orq experiment's rows instead of a local file. Requires ORQ_API_KEY. |
--experiment-run-id | str \| None | Specific run of --experiment-id to load. Latest run if omitted. |
--from-run | str \| None | Replay a previous run from .evaluatorq/sim-runs/: pass its file name, run id, path, or "latest". Re-runs the exact same personas, scenarios, and first messages; only the target/evaluators may differ. |
--memory-entity | str \| None / None | Memory entity_id sent with every agent:<key> (or bare <key>) target call, for agents with a memory store attached. Omit to mint a fresh id per conversation; pass one to reuse a specific (e.g. seeded) entity, shared across the run. |
--name / -n | str \| None / unset | Run name. Unset unless passed: the SDK then saves the run as sim and names an uploaded experiment simulation-<timestamp>-<id>. |
--sim-model | str \| None / unset | Model for the user-simulator, the judge, the recommendations pass and the executive summary. Unset, the simulated user, recommendations and executive summary run on the fast role (openai/gpt-6-luna by default) and the judge runs on the smart role (openai/gpt-6-sol by default); a value here applies to all of them. Sets llm_config.model. |
--max-turns | int / 10 | Maximum conversation turns. Unset unless passed; the SDK uses the replayed run's cap with --from-run, else 10. |
--datapoint-parallelism | int / 10 | Concurrent simulations. Unset unless passed; the SDK default is 10. --parallelism is a deprecated alias. |
--llm-parallelism | int / 10 | Ceiling on in-flight LLM requests for the whole run. -1 disables it. |
--evaluator | str (repeatable) / API defaults | Evaluator name(s). Repeatable. |
--save / --no-save | bool / True | Write the run to .evaluatorq/sim-runs/, with a run manifest. The SDK saves it after the recommendations and executive summary are attached. --no-save skips it. |
--recommendations / --no-recommendations | bool / True | Generate LLM remediation suggestions for failures, tied to their concrete cause. On by default; --no-recommendations skips the extra LLM call. Uses --sim-model. |
--results / -r | Path \| None / None | Path to write results JSONL. |
--report | Path \| None / None | Path to write full SimulationRun report JSON. The output reference describes this file and the --results rows. |
--report-md | Path \| None / None | Directory for an auto-named Markdown report. |
--report-html | Path \| None / None | Directory for an auto-named HTML report. |
--executive-summary / --no-executive-summary | bool / True | Generate an LLM narrative executive summary in the report. The SDK generates it before it saves the run. |
--yes / -y | bool / False | Skip interactive confirmation prompt. |
--verbose / -v | count / 0 | Increase verbosity. |
--quiet / -q | bool / False | Suppress non-error output. |
--config | PATH \| - / None | JSON file of simulate() keyword arguments, or - to read it from stdin. See Driving the CLI from a config file; eq sim schema prints the accepted shape. |
--llm-config | JSON / None | LLMCallConfig for every simulation-side LLM call as a JSON object. Merged field by field into "llm_config" from --config. --sim-model wins over its "model" when passed. |
--json | bool / False | Print the final SimulationRun as JSON on stdout. Progress and messages go to stderr; exit codes are unchanged. |
Driving the CLI from a config file¶
--config on eq sim simulate and eq sim run takes the keyword arguments of the Python simulate() and generate_and_simulate() functions as one JSON object, under their public names (run_name, experiment_description, raise_on_execution_failure). That reaches every data-shaped parameter, including the ones without a flag: inline personas and scenarios or datapoints, scoring, target_agent_timeout_ms, max_tool_result_chars, per_simulation_timeout_s, upload_results, the full llm_config, and on sim run edge_case_percentage. Pass a path, or - to read it from stdin. YAML is not accepted.
Inline datapoints, or personas with scenarios, count as the input source of eq sim simulate, so the file below needs no --input:
{
"target": "agent:my-agent",
"personas": [
{
"name": "Impatient customer",
"patience": 0.2,
"assertiveness": 0.8,
"politeness": 0.4,
"technical_level": 0.3,
"communication_style": "terse",
"background": "Wants a refund today"
}
],
"scenarios": [
{
"name": "Refund",
"goal": "Get a full refund",
"criteria": [{"description": "Agent asks for the order number", "type": "must_happen"}]
}
],
"max_turns": 6,
"llm_config": {"model": "openai/gpt-5.6-luna", "temperature": 0.2},
"target_reasoning_effort": "low",
"per_simulation_timeout_s": 300
}
eq sim run takes the generation keywords the same way:
{
"target": "agent:my-agent",
"num_personas": 4,
"edge_case_percentage": 0.25,
"generation_instructions": "customers replying in German",
"llm_config": {"model": "openai/gpt-5.6-luna"}
}
A flag passed on the command line beats the file, and the file beats the CLI default. The command checks where a value came from, not what it is, so --max-turns 10 wins over "max_turns": 6 even though 10 is the default. A null in the file means "not set" and is accepted only on fields that allow None, such as run_name, max_turns or datapoint_parallelism; the SDK default then applies. A value nobody sets (no flag, no file entry) is never sent: max_turns, datapoint_parallelism and run_name reach the SDK only when you pass them, so --from-run can still restore the replayed cap. "save": false in the file has the effect of --no-save. --llm-config and the file's "llm_config" merge field by field, and --sim-model wins for model.
The target and the input source are each one choice, so a command-line flag replaces the file's whole choice rather than adding a second one. --target, --vercel-url or --openai-model drops the file's "target". On eq sim simulate, --input, --dataset-id, --experiment-id or --from-run drops the file's "datapoints", "personas", "scenarios", "dataset_id", "experiment_id", "experiment_run_id" and "previous_run". --experiment-run-id alone narrows the file's "experiment_id" instead. Two flags for the same choice on the command line are still rejected.
Unknown keys fail the command before anything runs, at every depth, so "num_personas" in a sim simulate file or "patiense" inside a persona is rejected with the field path. "target" takes only the agent:<key> and deployment:<key> string forms; a callable or AgentTarget, user_simulator, judge, hooks and generation_client stay Python-only.
With --json, stdout carries the SimulationRun and nothing else; the run-store save, the report files and every progress line still happen, on stderr.
eq sim schema¶
Print a JSON schema: what --config accepts, or what --json prints.
| Flag | Type / Default | Description |
|---|---|---|
--input / --output | bool / --input | --input prints the config file schema. --output prints the SimulationRun schema, which both commands share. |
--command | simulate \| run / simulate | Which command's config file --input describes: SimulateRunConfig for eq sim simulate, GenerateAndSimulateRunConfig for eq sim run. |
eq sim upload-dataset¶
Upload simulation datapoints to an Orq dataset, or append them to an existing dataset.
eq sim upload-dataset -i cases.jsonl -n "Support simulation set"
eq sim upload-dataset -i more.jsonl --dataset-id <id>
The first command prints the new dataset ID. Use that ID, not the display name, with eq sim simulate --dataset-id or the Python extend_from_dataset() function. Here, upload-dataset --dataset-id adds rows to the stored dataset; extend_from_dataset() generates new cases without changing the stored dataset. Persona and scenario objects are JSON-stringified because the Orq dataset API accepts scalar inputs values. The simulation reader restores them when the dataset is used with eq sim simulate --dataset-id.
| Flag | Type / Default | Description |
|---|---|---|
--input / -i | Path (required) | Raw sim generate JSONL or a --dataset-format JSONL file. |
--name / -n | str \| None | Display name for a new dataset; required unless --dataset-id is provided. |
--path | str / Default | Orq folder path for a new dataset. |
--dataset-id | str \| None | Append to this existing dataset instead of creating one. |
eq sim generate¶
Generate simulation datapoints only — no simulation is run.
| Flag | Type / Default | Description |
|---|---|---|
--datapoints / -d | Path (required) | Path to write generated datapoints JSONL. |
--agent-description | str \| None / None | Free-text description of the agent. |
--target | str \| None / None | Agent target used to fetch the description when --agent-description is omitted. Accepts agent:<key>. |
--sim-model | str \| None / None | Model for persona/scenario/first-message generation. Unset resolves the fast role (openai/gpt-6-luna by default). |
--num-personas | int / 5 | Number of personas to generate. |
--num-scenarios | int / 5 | Number of scenarios to generate. |
--persona-seed | str (repeatable) / None | Archetype seed for a persona, e.g. "angry retiree" (repeatable). Each seed becomes one persona the LLM fleshes out — overrides --num-personas. Omit to auto-generate. |
--scenario-seed | str (repeatable) / None | Situation seed for a scenario, e.g. "disputes refund denial" (repeatable). Each seed becomes one scenario — overrides --num-scenarios. Omit to auto-generate. |
--generation-instructions | str / '' | Free-text steer applied to every generated persona AND scenario, e.g. "enterprise B2B buyers, replying in German". Stacks on top of any seeds and --edge-case-percentage. Empty leaves the built-in prompts unchanged. |
--dataset-format | bool / False | Write Orq dataset-row envelopes instead of raw simulation datapoints. |
--verbose / -v | count / 0 | Increase verbosity. |
--quiet / -q | bool / False | Suppress non-error output. |
eq sim from-traces¶
Build simulation datapoints from Orq production traces.
eq sim from-traces --output dp.jsonl --limit 50 --lookback-hours 24
eq sim from-traces --output dp.jsonl --extend 20 --agent-description "..."
Fetches recent traces from the Orq traces API (requires ORQ_API_KEY) and builds one datapoint per trace conversation: persona and scenario are inferred from the transcript, and the first message is the real user's opening message verbatim. Pass --extend N to additionally generate N new datapoints matching the traffic distribution of the fetched traces (extra LLM calls). Feed the output file to eq sim simulate --input to run it.
| Flag | Type / Default | Description |
|---|---|---|
--output / -o | Path (required) | Path to write generated datapoints JSONL. |
--limit | int / 20 | Maximum number of traces to fetch. |
--lookback-hours | float \| None | Only fetch traces from the last N hours. Default: no time filter. |
--search | str \| None | Free-text search applied to the trace list. |
--extend | int / 0 | Also generate N distribution-matched datapoints on top of the direct per-trace ones (extra LLM calls). 0 disables extension. |
--agent-description | str \| None | Agent description used for --extend generation. Optional; inferred from the traffic profile if omitted. |
--sim-model | str \| None / None | Model for persona/scenario inference and extension generation. Unset resolves the fast role (openai/gpt-6-luna by default). |
--llm-parallelism | int / 10 | Ceiling on in-flight LLM requests while building datapoints. -1 disables it. |
--verbose / -v | count / 0 | Increase verbosity. |
--quiet / -q | bool / False | Suppress non-error output. |
eq sim export¶
Export simulation results: OpenResponses payload JSON, or an HTML/Markdown report.
eq sim export --input results.jsonl --output payload.json
eq sim export --input sim-report.json --output report.html --format html --recommendations
Markdown/HTML exports include remediation suggestions if the input run JSON already carries them (runs generate them by default — see --no-recommendations), or if --recommendations is passed here to generate them at export time for a run that has none stored.
| Flag | Type / Default | Description |
|---|---|---|
--input / -i | Path (required) | Path to a results JSONL file or a SimulationRun report JSON (--report / --report-output). |
--output / -o | Path (required) | Path to write the exported file. |
--format | str / openresponses | Export format: openresponses (payload JSON), md (Markdown report), html (HTML report). |
--recommendations | bool / False | For md/html: generate LLM remediation suggestions at export time if none are stored. Extra LLM cost; uses --sim-model. |
--sim-model | str \| None / None | Model for --recommendations generation. Unset resolves the fast role (openai/gpt-6-luna by default). |
--target-label | str / agent | Target name shown in md/html report headers. |
eq sim validate¶
Validate a simulation datapoints JSONL file.
| Flag / Argument | Type / Default | Description |
|---|---|---|
PATH | Path \| None | Path to datapoints JSONL file (legacy positional form). |
--input / -i | Path \| None | Path to datapoints JSONL file to validate. (validate-dataset is retained as a compatibility alias.) |
eq sim validate-dataset (compatibility alias)¶
Deprecated alias for eq sim validate --input PATH. Retained for compatibility.
| Argument | Type / Default | Description |
|---|---|---|
PATH | Path (required) | Path to datapoints JSONL file to validate. |
eq sim runs¶
List recent simulation runs.
| Flag / Argument | Type / Default | Description |
|---|---|---|
DIRECTORY | Path \| None / None | Directory to scan. Defaults to .evaluatorq/sim-runs/. |
--limit / -n | int / 20 | Maximum number of runs to show. |
--full / -f | bool / False | Render at full content width; do not truncate columns. |
--json | bool / False | Emit runs as a JSON array on stdout. |