Changelog¶
All notable changes to evaluatorq are documented here.
[Unreleased]¶
- Four model roles,
fast,smart,classifierandembedding, name which model does which kind of work, with global flags to set them.eqtakes--fast-model,--smart-model,--classifier-model,--embedding-modeland a repeatable--model-override TASK=MODELahead of the subcommand;EVALUATORQ_FAST_MODEL,EVALUATORQ_SMART_MODELandEVALUATORQ_EMBEDDING_MODELjoinEVALUATORQ_CLASSIFIER_MODEL; the settings file gainsfast_model,smart_model,classifier_model,embedding_modelandmodel_overrides; and dashboard Settings shows one field per role. An unknown--model-overridetask exits with the list of valid tasks. A settings file saved before this change still loads: a non-defaultcompiler_modelorapply_modelbecomes a task override, and a value equal to the old default is dropped. The order of precedence and the task list are in Configuration › Models. - Deprecated:
EVALUATORQ_COMPILER_MODELandEVALUATORQ_APPLY_MODEL. Both still pin their task (finder.compilerandapply) and log one warning per process. Use--model-override,model_overridesin the settings file, or the role variables. - The Insights run form lets you choose the summary, classifier, embedding and (for the Question source) question compiler models from the Orq catalogue. A run saves the compiler model in its config so Re-run prefills it, and starting a run is rejected when the catalogue says the classifier cannot serve
/classifyor the embedding model is not an embedding model; with no catalogue the typed value is accepted and a warning is logged. - The Insights new-run page, the run-page "+ New run" dialog and "Re-run" now show one server-rendered three-step form (Traces, Analysis, Review). It combines both earlier forms: Presets, every source and label either form offered, individual coding-agent question toggles, custom questions, and Browse uploads for Finder exports and snapshots. A rejected submission re-renders with the entered values and the error, and review pages for saved runs now show the sign-in warning.
- The form estimates traces, cost and time before the run starts. The trace count is exact for a Finder export or snapshot and an
up toceiling for Recent and Question, taken from the Orq per-value counts and the trace limit. Cost multiplies traces by capped tokens per stage by the stage model's catalogue price; time uses the median per-trace seconds of your earlier Insights runs. Anything without a basis readsunknownwith the reason, and a total that leaves stages out sayspriced stages onlyortimed stages only. - Filter menus on Traces, Trace search and the Insights form show a trace count beside each value. Counts are Orq's for the search window, or a tally of the loaded rows on Traces, and values sort from most to least frequent. All three pages render the menu and chips through one
facet_pickermodule. - Built-in Top 10% views on /traces for slowest, costliest, token-heavy and largest-context traces and for costliest, token-heavy and longest conversations (BOPS-1243).
- New search on Traces reloads the table with its searched traces, so AI matches are visible in the table and related views.
- Ask AI can ask up to three questions of each trace, or none. The compiler now produces zero to three named classifier dimensions. A trace is included only when it matches all of them, and each dimension gets its own column in the Traces table, the Trace search table, and
eq findoutput. All of a trace's dimensions share one classify call, and a question that filters alone answer (such asAny traces above 50k tokens) runs no per-trace classifier at all. Within results now applies the question's metadata filters and token or duration bounds to the loaded rows. The run export moves toschema_version: 2: adimensionslist replaces the top-leveltaskandselection, and each trace carries per-dimensionanswersin place ofvalue,confidenceandprobabilities. - Ask AI progress is now visible while traces load and classify. The
/tracesprogress line shows live load counts, a thin progress bar appears for loading and classification, and the corner status badge is hidden while a run is active. - Within results questions set the table's filters. A bound or facet in the question (such as
over 20k tokens) narrows the Traces table to the matching loaded rows and shows up as a filter chip, in the Filters count and in the Filters menu; when no loaded row qualifies the run completes with zero rows and a notice instead of failing. Judging is capped by the AI trace limit in settings rather than by the number of rows in the table, AI answers stay on their rows when the table reloads until Clear AI results, and a wide table scrolls sideways inside the card with a pinned header and time column. - Ask AI only counts the answers a question asks for. The compiler no longer selects every choice label (for example
Neither) or the healthy label for a negated question; such a plan is rejected and retried once, so the match count reflects the traces the question is looking for. - Trace finder planning and filtering now explain more of what happened. Descriptive phrases such as
coding agentsbecome classifier dimensions instead of being treated as numeric-only questions, boolean selections accepttrueandfalselabels, model or provider filters can find model-level traces, and Within results names the filter and nearest loaded value when every row is dropped. Numeric-only plans with uncovered words show a warning that names the uncovered question text.
Notable defaults¶
- Red-team attack generation, all judges and evaluators, and apply-recommendations now default to the smart role,
openai/gpt-6-sol(wasopenai/gpt-5.6-luna). This covers red-team attackers and evaluators, the simulation judge, the defaultllm_jury()panel, the dashboard's apply merge and Insights summaries. It raises the cost of red-team, simulation-judge and apply runs. SetEVALUATORQ_SMART_MODEL=openai/gpt-5.6-lunato keep the old model. The simulated user, simulation generators and the trace-finder compiler use the fast role,openai/gpt-6-luna(wasopenai/gpt-5.6-luna), andDEFAULT_PIPELINE_MODELexports the fast role default. A call that passesmodel=explicitly is unchanged. With onlyOPENAI_API_KEYset, setEVALUATORQ_FAST_MODEL=gpt-6-lunaandEVALUATORQ_SMART_MODEL=gpt-6-sol, since the built-in ids are provider-prefixed. simulate()andgenerate_and_simulate()now generate remediation recommendations by default (recommendations=True), matchingred_team()and the CLI. That is one extra LLM pass per run over the results with a fixable failure. The returnedSimulationResultlist cannot carry them, so a call withoutsave=Trueorreport_path=generates them, discards them and logs a warning; passrecommendations=Falseto skip the call.- The minimum supported pydantic is now 2.12.
--configand--llm-configvalidation relies on per-callextra='forbid'(model_validate(..., extra='forbid')), which first shipped in 2.12, to reject misspelt keys at every depth, including inside models that ignore extra keys by default. eq sim simulateandeq sim runleave--name,--max-turnsand--datapoint-parallelismunset unless passed, so the SDK's defaults apply. The effective turn cap and concurrency are unchanged (10 and 10). Without--name, an uploaded experiment takes the SDK's generatedsimulation-<timestamp>-<id>name instead ofsim; the saved run is still namedsim. The CLI now saves through the SDK'ssave=, so a saved CLI run also gets a run manifest, and--save/--no-savereplaces the lone--no-save(which still works).eq redteam rungains--generate-strategies/--no-generate-strategiesand--cleanup-memory/--no-cleanup-memorypairs; the existing--no-*spellings still work. Every flag that has a--configfield now takes its default from the SDK, through one shared constant where the SDK resolvesNone(DEFAULT_DATAPOINT_PARALLELISM).- The Insights run form no longer has typed server-path inputs for Finder exports and snapshots; use Browse. The dashboard accepts Finder exports up to 10 MiB and snapshots up to 100 MiB; larger files go through
eq insights --from-finder PATHor--from-snapshot PATH. The single "Analyze coding agents" checkbox is replaced by per-question toggles, and a dashboard run now uses the models chosen in the form rather than only the defaults. - Balanced Trio and Single-Provider Trio replace GPT-5.6 Luna with
openai/gpt-6-lunaatmedium. The successor is both stronger on the same Intelligence Index v4.3.2 and cheaper at captured $0.10/$0.50 input/output rates per million tokens. Their flat estimated costs fall to $16.36 and $57.90 per 1,000 items respectively; the EU-hosted Luna seat is unchanged. - Open-Weight / Portable now seats
wafer/Kimi-K3at the host's default effort,none, instead of Baseten atmax. This accepted host-and-effort trade lowers the flat estimated panel cost from $37.66 to $34.28 per 1,000 items, using Wafer's captured $3.00/$12.75 input/output rates per million tokens. It does not establish equal benchmark accuracy at the two reasoning settings. - The Traces page (
/traces) automatically loads up to 200 traces from the last seven days with no user-entered filters and no AI calls; a saved project selection still applies. Press Load to fetch again with the selected controls. The table supports up to 5000 rows, and its default columns are Status, Started, Agent, Model, Tokens in, Tokens out, Cache read %, Cost, and Duration; the selected columns are saved in the dashboard settings file (default.evaluatorq/dashboard-settings.json, configurable withEVALUATORQ_DASHBOARD_SETTINGS) underexplorer_columns. - Ask AI on Traces defaults to "Just proceed" and, when traces are already loaded, to the "Within results" scope, which classifies the loaded population instead of starting a fresh search. Choose Review first under Ask AI on traces in Settings to see the plan before any classifier call is billed; asking within more than 500 loaded traces is forced through review regardless of that setting.
- Trace search remains available as a separate dashboard page at
/find; the new Traces explorer is at/traces. Both pages share the facet menu, filter chips, and trace drawer while showing tables suited to their result types. ModelInfocosts can now beNone. The catalogue used to drop any model Orq lists without a usable USD price; it now keeps the entry withinput_cost_per_1kandoutput_cost_per_1kset toNone, so the dashboard Settings model picker shows it. Calls to such a model stay unpriced, as before. Code that does arithmetic on the result ofget_model_info()must check forNone.ModelInfoalso gains a trailingmodel_typefield.- New
evaluatorq.formatspackage —ChatConversation,ResponsesConversation,OtelTraceandAtifTrajectoryconvert between any two formats via ATIF; ATIF is read at v1.7 and v1.8 (older versions raise;AtifTrajectory.from_jsonalso rejects a document with noschema_version, while a trajectory built in code defaults to v1.7) and written with the trajectory's own version (v1.7 for converter output), upgraded to v1.8 when a trajectory holds audio. - Jury presets run at each judge's default reasoning effort, and published costs assume output tokens equal to input.
JuryPreset.seated_efforts()now reports the catalog default for each seat (mostlymedium), which is the effort a preset call actually runs at, since presets send none; seats are chosen in this package at that effort rather than at each model's top effort. Publishedestimated_cost_per_1know assumes a flat 1,500 input and 1,500 output tokens per seat rather than 1,500 and 150, so every figure is about four times higher: Balanced Trio 17.56, Strong Jury 85.50, Open-Weight / Portable 37.66, EU Region 27.75, Single-Provider Trio 59.10.JuryPreset.priced_below_seated_effort()is removed, and the rate table'sseated_effortandpriced_at_ceilingfields are replaced bydefault_reasoning_effort. Strong Jury seatsanthropic/claude-opus-5-5in place ofanthropic/claude-opus-5, and Balanced Trio's reserve isdeepseek/deepseek-flash, the id that now serves DeepSeek V4.1 Flash.ModelInfogainsdefault_reasoning_effort. - New
evaluatorq.backends.CodingAgentTargetruns Claude Code, Codex CLI or OpenCode as the system under test, directly or throughorq launch, on the host or in a reusable per-clone container. Coding-agent turns now have a 300 s idle limit under a 2 h hard cap in host mode too; a host-mode agent that keeps printing can run for up to 2 h, where the previous 210 s effective ceiling cut it off. Container runs default to each agent's bypass mode and block Linux privilege escalation. Failures surface ascli.*error codes, including non-retryable timeout and container startup errors.SimulationRunnerpasses the target's ownmap_errorto the retry helper, so backend-specific codes reach simulation results where the default mapping used to apply. - A static red-team row with no user message now fails instead of being sent. The plain static pipeline used to send an empty prompt to an
AgentTargetoragent:target and score its reply, so a malformed dataset row could come back RESISTANT. It now raises before any target is created and appears under Failed Jobs, which the hybrid pipeline's static leg already did. CodingAgentTargetnow runs on a Windows host. Container mode usedos.getuid()andSIGHUP, the orphan sweep probed liveness withos.kill(pid, 0)(a Ctrl+C on Windows), and killing a timed-out turn usedos.killpg(); the first container turn raisedAttributeError, a leftover container made every start raiseOSError, and a host-mode timeout left the agent running. On Windows the container now runs as the image'sagentuser (uid 1001), the sweep asks the kernel whether the owner is alive, and a timed-out turn ends its process tree withtaskkill /T /F, best effort. Container cleanup there relies onatexitand the lease, since SIGTERM and SIGHUP are not delivered. The agent binary is now resolved on thePATHfromenv, so an npm-installedclaude.cmdis found on Windows; a.cmdor.batlauncher is refused withcli.unsafe_shimwhen an argument carries characterscmd.exewould interpret.- Insights labels and summaries now read a compact view of the conversation instead of the trace projection. The old projection kept the end of long traces within its byte budget and dropped the first user message in 82 of 100 sampled coding-agent traces; the new view keeps every opening user message, the start and end of each assistant message, and one line per group of tool calls, and leaves tool outputs out. Labels and summaries on long traces can change.
- The built-in
user_frustrationlabel now rates 1 (calm) to 5 (angry or giving up) instead of six levels from 0, so saved scores are not comparable with earlier runs. - Insights summaries no longer produce
tools_used; each trace instead records tool, shell-program, and skill counts taken from its messages astool_stats. eq insights --codingandinsights(coding_analysis=True)add coding-agent labels (task type, outcome, verification, scope creep, user corrections, unfixed errors, risky actions) to traces the classifier identifies as coding agents. Off by default.- Insights
--classifier-modelnow followsEVALUATORQ_CLASSIFIER_MODELor the dashboard setting when either is configured;typesafe/jev-latestremains the fallback. - Trace projections used by finder classification include each paired tool result's completion status and a fixed diagnostic category for recognized failures. Raw tool result bodies are omitted because they may contain credentials that pattern-based redaction cannot reliably identify, so classifications on tool-heavy traces can change. Insights allows more output tokens for summaries and cluster names after a live tool-heavy run exhausted the previous limits.
eq find --positive-onlykeeps only matching trace records in its JSON export. The terminal table still shows matches and a full-run summary, JSON counts describe the full run, and debug diagnostics still show every classifier call.- Trace Insights adds
insights(),eq insights, and a read-only dashboard page for analyzing saved Orq trace populations. Runs keep filtered population selection separate from fixed classifier labels and clusters discovered from summary text. It defaults toopenai/gpt-6-lunafor summaries and stores per-item summaries and embeddings in.evaluatorq/cache/insights.sqlite; installevaluatorq[insights]for the numerical dependencies and usecache=Falseor--no-cacheto bypass the local cache. Theconcerningpreset uses classifier level indices 0–4. In the CLI, the priority matrix usesintentby default when that dimension is selected, or the first selected dimension otherwise. The dashboard accepts Finder exports up to 10 MiB when starting an Insights run. - Saved Insights runs now open in a four-view review page by default. Themes, Activity, Map, Compare, and the trace list share filters that the URL restores. Activity reports recorded tool, shell-command, and skill counts, and distinguishes traces with missing activity data from traces with no recorded use. The dashboard Map is a display-only 2D projection of saved coordinates; it does not rerun embeddings or change the run. The new-run sheet supports validated Finder and snapshot uploads up to 100 MiB, individual coding-label selections behind the
coding_agentgate, and validated custom questions. Existing saved runs, exports, trace details, and legacy tab URLs remain available. - Insights runs save recorded token usage and known provider cost for tracked label, summary, describe, merge, and embed stages; semantic-query population compilation and filter-selection calls are excluded, and failed calls may have unknown billing. Missing or unpriced usage and excluded calls mean displayed dollars can be below whole-run spend. The dashboard shows known cost and partial price coverage in the run header, plus live trace counts for label and summary progress; a partial dollar amount is a lower bound.
eq dashboardnow reads.envfrom the launch directory and opens the local dashboard in a browser when the server starts. Already exported environment variables take precedence; the browser opens once, not on each hot reload.- Dashboard Settings can save an Orq profile, workspace, and project together. The selected profile supplies the API key and host, its workspace slug supplies Orq trace links, and the selected project ID limits dashboard Trace search. Changing profiles refreshes the available workspace and projects before saving. The CLI's masked profile keys are resolved from its private local credential file; no key is written to dashboard settings. The
eq findCLI uses the saved profile and project by default;--profileand--projectoverride them for one run. - The trace finder defaults to a seven-day window, a 500-trace population default (maximum 5000), 100-way classify concurrency,
openai/gpt-5.6-lunafor compilation, andtypesafe/jev-latestfor classification. Override the models through the dashboard Settings page; set the window, limit and parallelism per run on Trace search, through the finder flags, or through the documented environment variables. eq findnow shows a terminal activity indicator without printing every poll. Pass--debugto print changed progress and the compiler and classifier requests and responses, including trace content.inferenceno longer has to be passed with a recorded-output source. It now defaults to unset and resolves fromdata:ExperimentInputandTraceInputare replay sources by construction and resolve toinference=False, anything else toTrue.evaluatorq('run', data=TraceInput(limit=50), evaluators=[...])is now the idiomatic call; it used to raise, and the only way to learn the required flag was to trip that error. An explicitinference=Truewith a replay source still raises, andinference=Falseeverywhere it already appears keeps working unchanged.TraceInput.limitis nowint | None, defaulting toNone, and cannot be combined withtrace_id=. Trace mode never readlimit, soTraceInput(trace_id='t1', limit=50)used to validate and then return exactly one trace with no warning; it is now rejected. The resolved count (still 20 when unset) is thequery_limitproperty — read that rather than the field.- A trace query that selects nothing is no longer a green run.
fetch_traceswarns when a query matches no traces,evaluatorq(data=TraceInput(...))raises rather than evaluating zero rows, andred_team(datapoints=[])is rejected instead of running zero attacks and exiting 0. A mistypedsearch=used to produce a successful-looking report with a resistance rate computed over nothing Simulation matches them:simulate()/datapoints_from_traces()raise when aTraceInputyields no conversation with a user turn, where they used to return an empty list and report a run of zero personas. - One unreadable trace no longer kills the batch.
redteam.datapoints_from_tracesskips a trace that failed to import (or has no seedable turn) with a warning and keeps the rest, matching what simulation already did and whatfetch_tracesalways documented; it raises only when no usable trace remains. Failed imports still become individually-failing rows on the coreevaluatorq()path. All three surfaces now partition through the sharedcommon.trace_input.partition_traces. - The trace import now retries rate limits and server errors. The Orq trace endpoints have no SDK method, so
fetch_tracescalls them over rawhttpx;with_retryrecognised only the OpenAI SDK's own error classes, so a 429 or a 503 raised byraise_for_status()ended the import on the first attempt while the docstring promised retry.httpx.HTTPStatusErroris now retried on the same status rule as every other call path. - A naive
TraceInput.start_time/end_timeis now read as UTC, not as the host's local time. The Orq query fields are epoch milliseconds, so the sameTraceInputused to select a different window on every machine that ran it, and a naive bound compared against a timezone-aware one raisedTypeErrorinside model validation. Both fields are stored timezone-aware after validation; pass an aware datetime to select any other zone. orq_evaluator's result name now carries the model override.orq:<evaluator_id>whenmodelis omitted,orq:<evaluator_id>@<model>when it is set. Twoorq_evaluatorcalls differing only bymodel— the standard judge comparison — used to produce two result streams under one indistinguishable name. The evaluator also resolves its Orq client once per factory call instead of once per row, and acceptsapi_key=andbase_url=.- A recorded turn that both answered and called a tool now renders both.
output_to_textused to return the assistant text alone for anAgentResponse, so a judge asked whether the agent looked something up saw only the prose. Text and tool-call markers now render together throughmessages_to_text; scores from before this change are not comparable for mixed turns. - Classify answers now survive per repetition and appear in pairwise reports. The additive
JuryRepetition.raw_outputfield preserves the complete validated answer inEvaluationResult.raw_output["jury"]; pairwise observations retain each ordering's answer in saved JSON, while expanded HTML/dashboard comparisons show choice, confidence and probabilities in the run's original A/B frame. Missing details remain absent, and older saved records still load.send_results_to_orq()still stripsraw_output, so the hosted Orq experiment view does not receive these details. - Breaking for custom simulation hooks:
on_evaluator_complete(datapoint_id, name, score: float, result)is replaced byon_evaluator_complete(datapoint_id, name, score: EvaluatorScore, result: SimulationResult). Update the parameter names and read the verdict fromscore.score.value, its explanation fromscore.score.explanation, and failures fromscore.error. The hook now fires for errored and non-numeric evaluator outcomes too; numeric values alone entermetadata['evaluator_scores'], unusable outcomes entermetadata['evaluator_errors'], and reports count them asDropped. - Every judged
llm_jury()result now carries the full jury record onEvaluationResult.raw_output["jury"], on everyassignment—raw_outputwasNonethere before. The panel already computed a completeJuryResult(per-judge votes with model ID, verdict, explanation, abstain/failure state, and each judge's raw per-repetition verdicts), and the scorer attached it only underassignment="cyclic"; under the defaultassignment="all"the whole structure was discarded and the caller kept a one-line summary onexplanationnaming no judge and no per-judge verdict. Detail that was computed and billed cannot be recovered without re-running the panel, so it is now attached unconditionally rather than behind an opt-in a caller has to know to pass.JuryVote.repetitionsis now a list ofJuryRepetitionobjects, each carrying that pass'svalueandexplanation; saved reports with the old bare scalar list still load, with missing explanations asNone. A caller readingif result.raw_output is Noneoff a jury evaluator as a signal now takes the other branch — nothing in evaluatorq did, andsend_resultsstill stripsraw_outputbefore uploading to the Orq platform, so this is local-only data either way. The one case that still carries no record is a datapoint whose target errored: no panel ran, so there is nothing to record, and it returnsinconclusivewithraw_output=Noneas before. - The jury summary line omits the agreement clause instead of printing
raw agreement n/a. A single cyclic vote and a numeric panel both have no cross-judge agreement to report, and a placeholder sitting in a fixed slot reads as a measurement that came back empty rather than one that does not apply.[jury: 3/3 judges, raw agreement 67%]is unchanged;[jury: 1/1 judges, raw agreement n/a]is now[jury: 1/1 judges]. Update any assertion or log scraper matching the literalraw agreement n/a. - A panel that ran but reached no verdict now records its judge failure under
EvaluationResult.raw_output["evaluation_error"]. The value uses the sameRunErrorshape as red-team evaluation errors and includes the last judge error, all recorded judge errors, and the failed-judge count. The key is absent when judges abstained cleanly or simply failed to reach the quorum without recording an error. A target error still leavesraw_output=Nonebecause no panel ran, and the directllm_jury()path leaves the payload inraw_outputfor the caller to inspect; only the red-team report converter promotes it to the typed run error rollup. - Pairwise
JudgeStats.consistencyandconsistency_raware now populated under bothpluralityandbt-sigma. They measure repeated-pass self-consistency directly from the observations and do not require a Bradley-Terry fit;sigmaremains available only from the BT-sigma fit. HTML reports show consistency columns when any judge has data, usen/afor judges without measurable repeats, and explain the empty case in a caption. to_open_responses()now preserves simulation state in fixed locations. Its metadata always includescriteria_verified,criteria_meta,criteria_errors,scorer_errors, anddatapoint_id, and the response envelope includestoken_usage_knownbesideusage. On the evaluatorq job path,criteria_errorsandscorer_errorsareNonebecause conversion happens before scoring; an unserialisable metadata value is published as itsrepr()with a warning.- Built-in simulation evaluator detail now has structured local output.
criteria_metattaches its criteria records underraw_output['criteria']with anauditedlabel, andconversation_qualityreturns aConversationQualityScorewhosebreakdowncontains component scores and weights. Thisraw_outputis available on the returnedEvaluationResultand local result dumps, but is stripped before the Orq experiment upload and is not attached to the evaluator span. - Returned simulation results now retain evaluator detail. Each
SimulationResult.evaluator_detailsmapping contains the structuredraw_outputfor evaluators that provide it, including criteria evidence forcriteria_met, component scores and weights forconversation_quality, and custom evaluator output. The field is also preserved in savedSimulationRunJSON and JSONL result exports; the mapping is empty when no evaluator provides structured output. on_stage_end()now receives stage failures inmeta['error']as the live exception object, including forgenerate().DefaultHookslogs the failure atWARNINGlevel andRichHooksprints a failed-stage line; custom hooks should test whether the key isNonebefore formatting the exception.sim_model=is removed from every simulation entry point;llm_config=carries the model.simulate(),generate_and_simulate(),generate(),generate_personas()/generate_scenarios()(and their singular forms) andextend_from_experiment()no longer accept the keyword — passllm_config=LLMCallConfig(model=...), which says the same thing and can carrytemperature,reasoning_effort,timeout_ms,extra_bodyand a client beside it. Two spellings of one setting could not report an explicitly-passed default as a contradiction:simulate(sim_model=DEFAULT_MODEL, llm_config=LLMCallConfig(model='other'))ran onotherand warned about nothing. Update any call passingsim_model=— there is no deprecation shim. Unaffected: the CLI's--sim-modelflag, which now builds that config for you, and themodel=argument onSimulationRunner, the generators and the trace helpers, which still folds into a config beside it.JudgeAgentpins itself to the Responses API and logs when a config says otherwise.llm_config.apiis one knob for both simulation agents, and only one of them can honour every value: the judge sends function tools andreasoning_effortin one request, which chat completions answers with a 400 on models likegpt-5.4-mini, while the user simulator's plain completion works on either endpoint. Same model, two roles, two viable protocols.LLMCallConfig(api='chat_completions')now applies to the user simulator and is overridden on the judge with aWARNINGnaming it, where it previously reached the judge and broke the run. The pin isJudgeAgent.REQUIRED_API; the general default is stillBaseAgent.DEFAULT_API.- Every simulation agent defaults to the Responses API when its
LLMCallConfigleavesapiunset, overriding that class's ownchat_completionsdefault.JudgeAgentConfigandUserSimulatorAgentConfigsupplied this before; a caller handing an agent a bareLLMCallConfigused to get chat completions instead, which returns a 400 on models likegpt-5.4-minithe moment the judge sends function tools andreasoning_efforttogether. The default isBaseAgent.DEFAULT_API— a class attribute, not an environment variable, because it is a per-call setting. SetLLMCallConfig(api='chat_completions')to opt out. - A job that returns a top-level
errorkey now marks the row as failed, instead of counting as a clean success.process_jobread onlynameandoutputfrom a job's return value, soJobResult.errorwas set only when the job raised. A job that caught its own failure and reported it — which is what the simulation jobs do, because one dead row must not kill the batch — came back witherror=None, and a run whose conversation never happened printedFailed Jobs 0,Success Rate 100%and exited 0. Both simulation jobs —simulate()'s own andwrap_simulation_agent's — now emiterrorunconditionally (Noneon success, the runner's reason when the run ended inerrorortimeout), andprocess_jobhonours the key while keeping the output, so the transcript survives for diagnosis — its evaluators are skipped, though, because scoring a conversation already known to be dead buys nothing and costs an LLM judge call per row.wrap_langchain_agentnow catches a failinginvoke/ainvokeand reports it the same way rather than letting it raise; a caller mistake (no prompt and no usablemessagescolumn) still raises. The job's OTel span is now markedERRORon every failed row, including the two raise paths, which previously closed asOK. Presence is the failure signal, not the message's truthiness: a payload that flattens to nothing is reported asjob reported a failure with no readable messageand logged, rather than silently becoming a clean row.check_pass_failures(results, treat_errors_as_failure=True)catches these rows as a result — that flag still defaults toFalse, so a caller who does not pass it gets the same exit code as before; the change is to the reported counts, not to any exit status. Any other job already returning a top-levelerrorkey for a non-failure reason will now be counted as failed. simulate(exit_on_failure=True)now raises when a run ended inerrorortimeout, not only when a row was dropped. The check counted rows missing from the result cache, and a run the runner ended inerroris in the cache — soeq simandsimulate()returned normally over a target that was 401 throughout, which is what the caller asked to exit on.SimulationDroppedError's message now names both counts.exit_on_failure=Falseis unchanged: the same rows are reported as aWARNING. Scorer verdicts are still reporting only — an agent that answered badly does not raise.LLMCallConfig.temperaturehas no default — unset means the parameter is not sent, and the provider applies its own. It previously defaulted to1.0, and evaluatorq's own call sites layered literals of their own on top (0.8for persona and first-message generation,0.9for edge-case scenarios,0.7for the executive summary and chat-completions agent calls,0.3for trace analysis,0.0for the judge). Reasoning-class models reject the parameter outright rather than clamping it —gpt-5.6-lunaanswers400 Unsupported parameter: 'temperature' is not supported with this model— so a hardcoded temperature on the default model turned every persona x scenario pair into a failure andsimulate()into aRuntimeError. No call site sends a temperature now; a caller who wants one setsLLMCallConfig(temperature=...), and a per-calltemperature=Nonemeans unset rather than an explicit null.BaseAgent._resolved_temperaturelost itsfallbackargument accordingly and now gates onmodel_fields_setlike its sibling resolvers, so an explicitLLMCallConfig(temperature=None)opts an agent out rather than deferring to a call site. The judge is the one behaviour change worth planning around: its scoring calls are no longer pinned totemperature=0.0, so judgments are no longer reproducible run-to-run by default. That affects everysimulate()caller, not only those on a reasoning model. PassJudgeAgentConfig(temperature=0.0)to restore it — on a model that accepts the parameter.- A set-but-empty
ORQ_OTEL_*tuning variable now logs aWARNINGand falls back to the default, instead of falling back silently._env_inttreated an empty string like an unset variable, so an unresolved workflow variable in a CIenv:block expands to the empty string and disabled the knob with no signal. Whitespace-only values are treated the same way. Related:ORQ_OTEL_MAX_BATCH_SIZElarger thanORQ_OTEL_MAX_QUEUE_SIZEis still clamped down to the queue size, but the clamp now announces itself with aWARNINGrather than happening silently. EVALUATORQ_REASONING_EFFORThas no default — unset means the parameter is not sent, and the model applies its own. It previously fell back to"medium"for the simulator's own calls (user simulator, judge). A global effort is the wrong default in both directions: on a model that does not accept the parameter it costs a rejected request plus a retry per(model, tool shape)— memoised per process, so a short run or CI job never amortises it — and on a model that does, it silently overrides the provider's own tuned value. Set the env var, orLLMCallConfig.reasoning_efforton the agent's config, when you actually want a specific effort. Simulation only; red teaming'starget_reasoning_effortwas already opt-in.OrqResponsesTarget.retry_attemptsnow defaults to1— a single attempt, no retry — down from falling through towith_retry's default of5.common.target_call.call_target_with_retryis the single retry owner for target calls on every surface that drives a target (red team static, hybrid, pipeline, orchestrator, and simulation); a target that also retries internally multiplies against that budget instead of adding to it — 5 inner attempts under 3 outer ones is 15 calls to a target that is already refusing. Raiseretry_attemptsonly when constructing the target directly and callingrespond()outsidecall_target_with_retry.- Env-var overrides now share one reader, and a misconfigured
EVALUATORQ_LLM_TIMEOUT_S/EVALUATORQ_LLM_MAX_TOKENSwarns and falls back to the default instead of raising.common.env_config(env_int/env_float/env_bool) is now the single place env overrides are parsed and validated: unset falls back to the default silently, and a set-but-empty/whitespace, unparseable, out-of-range, or non-finite value logs aWARNINGand falls back to the default. It never raises. The private readers intracing/setup.pyandsimulation/agents/base.pyroute through it. This changes the two simulation knobs above, which previously raised aValueErroron a non-numeric value and crashed the process at import (DEFAULT_MAX_TOKENSis computed at module scope); they now warn and use the default, matching the non-fatal contract theORQ_OTEL_*tracing knobs already followed, and both are now bounded withmin_value=1so0, a negative, ornan/infalso fall back with a warning rather than reaching the provider.ORQ_DISABLE_TRACINGalso now recognisesyes/onand is case-insensitive, in addition to the previous1/true. A shared reader withenv_int/env_floatexisted briefly on the RES-1286 branch and was removed before merge in favour of pydanticFieldbounds on the recommendations config (which reject a meaningless value instead of warning and falling back); this reintroduces it deliberately for the process-global tuning knobs, where a warn-and-continue contract is wanted over a hard failure.EVALUATORQ_CAPTURE_MESSAGE_CONTENTandEVALUATORQ_PROPAGATE_TRACE_CONTEXTroute throughenv_booltoo, so they now recogniseyes/on/no/offcase-insensitively and warn on an unrecognised value instead of silently reading it as false — previously anything buttrue/1turned the toggle off with no signal.EVALUATORQ_CATALOGUE_TIMEOUT_Sroutes throughenv_float(min_value=0.1)and no longer raises at import on a non-numeric value.EVALUATORQ_SPAN_MAX_TEXT_CHARS,EVALUATORQ_REASONING_EFFORT,ORQ_DEBUGandCOLUMNSare left as bespoke reads (a capture-all sentinel whose default isNonerather than a number, a string enum, a debug-only truthy-any toggle, and a terminal probe respectively). LLMCallConfig.completion_params()and.responses_params()are removed. Both are replaced by a singleLLMCallConfig.request_params(*, api=None, **params), which renders the shape theapiargument names (self.apiwhen omitted). Two builders meant a call site could render a config that saysresponsesinto chat-completions shape and no one would notice — exactly the accepted-then-ignored failure this class exists to prevent. A call site that is structurally single-endpoint passesapi=explicitly and gets a warning if that contradicts an explicitly-setself.api. Update any call site using either removed method — there is no deprecation shim.- Simulation's Responses calls and the executive-summary narrative now go through the canonical executors, and are priced.
BaseAgent._call_responsesroutes throughcommon.llm_call.execute_response, andgenerate_executive_summaryroutes throughexecute_chat_completion— both now get slot limiting, the reasoning drop-and-retry-once, pipeline metadata, trace headers and, previously missing, aprice_usagecall. Simulation Responses calls on non-Orq endpoints were previously left unpriced entirely.generate_executive_summarynow returns anExecutiveSummary(text, usage)dataclass instead ofstr | None— update any caller that unpacked or compared the old return value directly. -
Post-processing spend now lands in the run totals instead of only the log. Red team's
ReportSummarygainspost_processing_token_usage(recommendation generation, including trace condensing, plus the executive summary), folded intotoken_usage_totalonce both steps have run, been skipped, or failed — previously that spend was log-only and invisible toreport.summary. Simulation's newSimulationRun.token_usage_totalsums every result's usage plus, forgenerate_and_simulate(), the GENERATE stage's persona/scenario cost and the executive summary's cost; it does not include recommendation generation, which stays log-only on both surfaces. -
turn_efficiencyno longer pays more for a longer conversation. The scorer's cliffs stopped at 6 turns and the decay past them restarted from1.0, so a 7-turn run scored0.9while a 6-turn run scored0.7— the curve went up at the seam. The decay now continues from the last cliff's score (7 turns →0.6, 8 →0.5, floor0.3), so scores fall monotonically with turn count. Only runs that optedturn_efficiencyorconversation_qualityintoevaluator_namesare affected — neither is in the default["goal_achieved", "criteria_met"]. Both scorers' policy is now exposed assimulate(scoring=SimulationScoringConfig(...))/generate_and_simulate(scoring=...): the cliffs, the decay, the floor and the composite weights, all with the shipped values as defaults. The config rejects unordered cliffs and weights that do not sum to1.0at construction. -
Config that was accepted and silently ignored now takes effect. Several fields validated cleanly and then did nothing, so a caller who set them saw no error and no change. All of them are now honoured, which means runs change behaviour for callers who were already setting them: simulation's
BaseAgentread onlymodel/api/client/retry_countfrom itsLLMCallConfigand fell back to literals for everything else, sotemperature,max_tokens,timeout_ms,extra_kwargsandreasoning_efforton aJudgeAgentor user simulator were inert;LLMConfig.max_tool_continuationswas documented as user config but read from the module-levelPIPELINE_CONFIGdefault, so a tool-using ORQ target still gave up after 5 continuation rounds and reported "unresolved pending tool calls" — which the judge then scored as a genuine weak answer; and red team's static leg droppedtarget_agent_timeout_msandmax_target_retriesentirely (the model/deployment path ran on a hardcoded 300 s ceiling with no retry). Per-call-site literals are now defaults that a caller-supplied value beats, in that order. extra_bodyis a dedicated parameter, not anextra_kwargskey — and passing it inextra_kwargsnow raises. It carries the Orq router body (retry policy, thread and memory ids), which is owned by the call site, so a user routing it throughextra_kwargssilently replaced the router body rather than adding to it: retry hints vanished with no signal on the judge, jury and recommendation paths.execute_chat_completion,execute_chat_parse,execute_responseandgenerate_structuredall gainedextra_body=;check_reserved_keysrejects the key insideextra_kwargson both endpoints. Update any call site usingextra_kwargs={'extra_body': ...}— it now raisesValueErrorinstead of quietly winning. The one merge seam isgenerate_focus_area_recommendations, where a caller-suppliedextra_bodymerges into the router body.LLMCallConfig.extra_bodyis the user seam into the request body, and it merges rather than replaces. There are two injection points and they target different parts of the request:extra_kwargssets top-level SDK call arguments,extra_bodysets fields in the HTTP body the SDK has no named parameter for — which is where the Orq router's retry policy and thread/memory ids live. Previously only the call site could write that body, so a user had no way to add a body field at all; the one workaround,extra_kwargs={'extra_body': ...}, replaced the whole thing, silently dropping the router's retry hints. The new field is layered over the call-site body per key: your keys win, the ones you did not set survive. Honoured on both endpoints (LLMCallConfig.request_params(), which renders the shapeapinames), by the judge, the simulation agents and the Responses target.- One reserved-key vocabulary instead of three.
contracts.pyowns_RESERVED_COMPLETION_KEYS/_RESERVED_RESPONSES_KEYSand the sharedcheck_reserved_keys;common.structured_outputno longer keeps its own narrower copy, which had drifted to omitextra_bodyon the chat legs and to reservemax_output_tokenswhile leavingmax_completion_tokensoverridable. The guard also now runs inside thecommon.llm_callexecutors, so a caller that reaches one without anLLMCallConfiggets the same protection. reasoning_effortreaches the provider on every role, and on both endpoints for the target, judge and jury. It was read at exactly one call site (OrqResponsesTarget), so setting it on the attacker, the judge, the jury or a simulation agent validated and did nothing.LLMCallConfig.request_params()now renders the flat chat-completions spelling or the Responsesreasoning={'effort': ...}block depending on the endpoint (self.api, or an explicitapi=at a structurally single-endpoint call site). The attacker role is chat-completions only: its call sites passapi='chat_completions'explicitly, soattacker=LLMCallConfig(api='responses')does not switch the endpoint. The two sites that build params throughrequest_params()warn when you set it; the rest pass their fields to the executor directly and cannot warn.EvaluatorConfignow subclassesLLMCallConfigand inheritstemperature,max_tokens,timeout_ms,extra_kwargs,extra_body,reasoning_effortandclientverbatim — it previously hand-listed them and setextra='forbid', soLLMConfig(evaluator=LLMCallConfig(reasoning_effort=...)), the documented pattern, raised aValidationError. Three fields are still declared on the subclass deliberately:model(nullable judge shorthand, excluded from dumps),api(defaults toresponses, notchat_completions) andretry_count(identical declaration, judge-specific description). An import-time guardrail compares each inherited field's annotation, default and constraints against the parent and fails loudly on any other override — or on an allowlist entry whose divergence has disappeared; a plain field-name comparison could not see a default drifting apart, which is the failure that actually occurred.red_team()'starget_reasoning_effortreached one of five target constructions and its preflight could raiseValueErrorfor targets that never receive it; the preflight is now scoped to targets whose backend actually sends it, and the rest warn.simulate()andgenerate_and_simulate()gainedtarget_reasoning_effortso the two surfaces measure the same agent.- The judge's Chat Completions legs now forward
reasoning_effortandextra_bodytoo.run_judgepassed onlyextra_kwargson both of them, so an effort or a router body set on anEvaluatorConfig— or viallm_jury(reasoning_effort=...)/llm_jury_pairwise()/PairwiseComparator, which build their config internally — worked on the Responses judge and was silently inert the moment the call fell back to chat completions (structured_output=False, a client that does not route through the Orq router, or a model the Responses endpoint will not take). Jury verdicts change for callers already passingreasoning_effort=. Red team's recommendation call sites likewise never readEvaluatorConfig.extra_bodyand never passedreasoning_effortat all; the router body there is now layered call-site-first, so a caller key wins per key. - A failed model-catalogue fetch no longer poisons the process. One HTTP hiccup against
GET /v2/modelscached{}for the whole run, silently degrading every later call to unpriced and chat-completions-only. The fetch is now retried up to three times before the empty result is cached, and its timeout is settable viaEVALUATORQ_CATALOGUE_TIMEOUT_S. Newregister_model()/get_model_info()/clear_model_overrides()make the catalogue writable: a model Orq does not list — a self-hosted deployment, or one newer than the workspace catalogue — could previously neither be priced, qualified for Responses, nor have its reasoning effort validated, and there was no way to see or fix that. register_model()keys on the bare model id, andModelInfonormalizes an emptyreasoning_effortstoNone.'openai/gpt-x'and'gpt-x'name one model and now register one entry: theprovider/prefix is stripped on write, because every internal lookup asks for the bare spelling — a qualified registration stored verbatim was never found, so the override was a silent no-op. Registering both spellings replaces rather than duplicates.ModelInfo(reasoning_efforts=frozenset())becomesreasoning_efforts=None("the catalogue does not say"), the only other reading the field allows: an empty set madevalidate_reasoning_effortreject every value, defaults included, with an empty accepted-values list.- Red team's static leg reports the same error taxonomy as hybrid. It passed neither the run config nor the backend's
map_errorto its job factory, so an identical target failing identically producedorq.http.429in hybrid mode and a generictarget_errorin static. The self-judge / family-bias guard was inert there too — it compared raw target strings likeagent:abc123, which resolve to provider familyunknown, so judging gpt-5 with gpt-5 warned in dynamic mode and never in static. - Previously-hardcoded budgets are now parameters, with their current values as defaults:
with_retry's backoff curve (min_wait_s,max_wait_s,jitter_fraction),generate_structured's 300 s timeout, the executive summary's 400-token cap,apply_recommendations' four merge budgets, simulation's 500-char tool-result truncation (which shapes what the simulated user reacts to and the judge then scores), the adversarial-timeout abandon threshold, the objective-generation batch size, and the black-box probe turn budget. The executive summary also stopped raising a swallowedTypeErroronextra_kwargs={'temperature': 1}— the very escape hatch its own docstring documents for reasoning-class models — which surfaced as a silently absent summary. - New CLI flags.
eq redteam rungained--target-timeout-ms,--max-target-retries,--retry-count,--max-tool-continuationsand--target-reasoning-effort;eq sim rungained--target-reasoning-effort,--persona-seedand--scenario-seed.simulate()/generate_and_simulate()gainedtarget_agent_timeout_ms,max_target_retries,per_simulation_timeout_s,max_tool_result_charsandedge_case_percentage— the per-simulation wall clock previously existed only onrun_batch, which the public path bypasses, so a stalled conversation had no overall bound andTerminatedBy.timeoutwas unreachable. SimulationConfig's four new numeric knobs are bounded, andper_simulation_timeout_s=0now raises.target_agent_timeout_msandmax_tool_result_charsrequire> 0,max_target_retries>= 0, andper_simulation_timeout_s> 0—Noneis the only spelling of unbounded.0previously validated and then reached atimeout_s <= 0sentinel that read it as no bound at all, the opposite of what typing it means.EVALUATORQ_LLM_TIMEOUT_S,EVALUATORQ_LLM_MAX_TOKENSandEVALUATORQ_REASONING_EFFORTare resolved at call time, not import time. Settingos.environ[...]after importing evaluatorq was previously a no-op, and the values could not differ per agent or per run. They are now fallback defaults that a config value beats. Still simulation-only.-
The Orq Responses target honours the config it is handed and survives a model that rejects reasoning. It built its request from four fields and dropped
temperature,extra_kwargs,apiandretry_count, which meant no Responses option outside that set —top_p,store,truncation,tool_choice,parallel_tool_calls, … — was reachable at all. It also never applied the drop-and-retry-once contract every other Responses call uses, so a target model that 400s on thereasoningblock failed every datapoint in the run rather than retrying once without it. Reasoning effort is now recorded as its own span attribute, so two runs at different efforts are distinguishable in traces even withEVALUATORQ_CAPTURE_MESSAGE_CONTENT=false. -
Structured generation degrades through four chat rungs instead of two, and parses what it gets back.
generate_structurednow runs strictparse()→ non-strictjson_schema→ a forced tool call carrying the same schema → barejson_object, and every text rung's output goes through the canonical fence-tolerant parser (common.extract_json) plus schema validation before being returned. Callers that used to receive(None, raw_content)on any fallback and fence-strip it themselves now get(parsed_model, raw_content)— three call sites (persona_generator,scenario_generator,traces) never did that stripping and dropped a fenced``json payload on the floor.parsed is Nonenow uniformly means "no rung produced anything that validated", with the last non-empty text still returned to log. The forced-tool rung is what makes this a *stricter* backup rather than a looser one:tool_choicenaming the function leaves the model no prose channel, and function calling is a different provider capability thanresponse_format, so models that 400 on a JSON schema often still support it. **It is the only rung that edits the prompt** (one appended user turn telling the model the tool will be called) and it is skipped entirely when the caller passed their owntools/tool_choice. **This can cost up to two extra calls** on a provider that answers nothing usable — previously the ladder gave up after two. Truncation and refusals still raise on every rung rather than continuing. The rung that answered is recorded on the span asorq.structured_output.leg; the old booleanorq.structured_output.fallbackis still written. (Return shape: as of the usage change below, the pair is now theparsed/rawfields of aStructuredResult`.) generate_structuredreturns what the ladder cost, and every rung is counted. It returns aStructuredResult(parsed,raw,usage) instead of a(parsed, raw)tuple — unpacking the result no longer works; use the fields. The dataclass exists because the call can bill up to five provider requests (the Responses leg plus four chat rungs) and none of that usage was extracted at all, so a call that degraded twice was reported as free.usageis the sum over every rung that reached the provider, not just the one that answered, priced per rung throughprice_usage. A rung whose usage block cannot be read is counted as one unpriced call with a warning naming the cause — never as zero — sopriced_calls < callsmarks the figure as a lower bound. The eleven call sites (persona/scenario generators,traces.py, and bothrecommendations.pymodules) all sit on paths with no report field for this spend, so each phase logs its own total (Persona/scenario generation: N tokens over M LLM call(s), $X);PersonaGenerator/ScenarioGeneratorgainedget_usage()/reset_usage()mirroringBaseAgent. Private helperstraces._summarize_conversationandredteam.reports.recommendations._condense_attacknow return their usage alongside their result. (RES-1295)generate_structureddisarms the client's SDK retry budget before the ladder runs. Every rung is already wrapped inwith_retry, so the two layers multiplied — on evaluatorq's own default path, not just on an injected client:resolve_llm_clientbuilds clients withmax_retries=2and red team's recommendation client withretry_count=3, so five outer attempts over two SDK retries was fifteen requests per rung. It now callswithout_client_retriesonce at the top, ascommon.judgealready did. Callers who were relying on the stacking see strictly fewer requests on a 429/5xx storm — raiseretry_countto compensate. The clone never mutates the client you passed.- A structured-generation call that raises still reports what it billed. New
StructuredGenerationErrorincommon.structured_output(subclassingRuntimeError, so anexcept RuntimeErrorkeeps working) carries the ladder total on.usage; a provider error propagates as itself — masking a 429 behind aRuntimeErrorwould cost the caller the status code — with the same attribute attached in place, so callers harvest it withgetattr(exc, 'usage', None)rather than by exception type. Truncation is the common case:LengthFinishReasonErrorhands back the completion it refused to parse, so aparse()rung that generated a fullmax_tokensof output is no longer counted as free. The three legs that swallow the exception — trace summarizing, and bothrecommendations.pymodules — now fold that spend into their phase totals, which previously reported "no usage reported by the provider" for an entire failed ladder. - Judge verdicts are generated explanation-first, and a self-contradictory abstention is coerced.
EvaluatorResponsePayloadand the dynamic jury verdict model now declarevaluelast, so the structured-output schema makes the model write its reasoning before committing to a verdict rather than justifying one it already emitted;DEFAULT_SECURITY_EVALUATOR_SYSTEM_PROMPTwas reordered to match, since a prompt that listsvaluefirst fights its own schema. A payload that arrives withabstain=Trueand a non-nullvaluenow has the value dropped, with alogger.warning, ajudge.verdict_coercedspan attribute andJudgeOutcome.verdict_coerced. Previously that contradiction flowed on unchecked: the static OWASP and adaptive evaluator paths readpayload.valuewithout consultingabstainat all, so a judge that declined to answer was still scored as a decisive pass or fail, and any vote built straight from the pair would tripJuryVote's validator after the judge had been billed. The system prompt now also describesabstain— it previously named only two keys while the enforced schema had three. The mirror shape (value=None,abstain=False) is unchanged: the redteam paths still count it as a failed repetition,llm_jurystill reads it as an abstention. - The simulation pipeline runs its own LLM calls on the Responses API.
UserSimulatorAgentConfignow defaultsapi='responses'(the judge already did), and the persona, scenario and first-message generators call the Responses endpoint, so a run's spans are uniformlyresponses {model}instead of a mix the trace UI types differently — which is what made the simulator and generator calls look absent when scanning for router spans. Passapi='chat_completions'on the config to opt out.generate_structuredgainsapi=with the default unchanged, so the red-team report paths are unaffected; onapi='responses'it degrades to the chat legs, with a warning naming the cause, when the endpoint is absent (404), when a 400's body names the schema form as unsupported, or when nothing comes back parsed. Any other 400 raises rather than being blamed on the provider. The user simulator no longer sees the target's tool calls or tool results — after role inversion those rows are ausermessage carryingtool_callsfollowed by orphantoolrows, which the provider rejects, so every simulation against a tool-using agent died at turn 1. The canonical transcript is untouched: the target and the judge still see the tool traffic. - Scenario and persona generation size their
max_tokensfrom the count they were asked for, instead of a flat cap. A batched structured call for ~15+ items truncated, and truncated structured output is unrecoverable (both legs raise rather than retry at the same budget). This raises spend for large counts — a 30-scenario request now asks for 15 000 output tokens where it asked for 6 000. The budget is deliberately unbounded at the top: a count large enough to exceed the model's own limit fails with the provider naming the limit, which beats silent truncation at an arbitrary ceiling. - One default model for every surface:
openai/gpt-5.6-luna.evaluatorq.contracts.DEFAULT_PIPELINE_MODELis now the single source;simulation.types.DEFAULT_MODEL,llm_jury.DEFAULT_JUDGE_MODELand the dashboard'sDEFAULT_APPLY_MODELare aliases of it rather than three independently drifting literals (previouslygpt-5-mini,openai/gpt-5.4-mini,openai/gpt-5.4-miniandgpt-5.6-luna). Each surface stays individually overridable —--attack-model/--evaluator-model,--sim-model,judges=/model=,EVALUATORQ_APPLY_MODEL— only the fallback is shared. Every default is now provider-prefixed for the Orq router; targeting OpenAI directly (onlyOPENAI_API_KEYset) means passing the baregpt-5.6-luna, which the red-team defaults did not previously require. DEFAULT_TARGET_MAX_TOKENSis10_000, up from5000, and simulation now shares it. On a reasoning model the budget covers hidden reasoning as well as the visible answer, so the old figure left much less room for output than it used to — truncation surfaces as a judgePARSEerror (incomplete_details.reason=max_output_tokens) or, in simulation,finish_reason=lengthbefore any tool call. It capsLLMConfig.max_tokens(attacker),EvaluatorConfig.max_tokens(judges),OpenAIModelTarget,create_deployment_job, and now simulation'sDEFAULT_MAX_TOKENS, whoseEVALUATORQ_LLM_MAX_TOKENSoverride is unchanged (its default was a separate8192).- The dashboard's apply-recommendations merge no longer has its own default model. It was deliberately set apart from the pipeline's — rewriting production agent instructions warrants a stronger model than scoring does — and now shares the fallback like every other surface.
EVALUATORQ_APPLY_MODELis how you raise it again, unchanged. - Multi-turn requests now carry Anthropic prompt-cache breakpoints, on Orq-routed clients only. Simulation agents and the red-team
OpenAIModelTargetmark the system prompt and the end of the persisted transcript withcache_control: {"type": "ephemeral"}— on the Chat Completions path as a marked content block, on the Responses path as a markedinputcontent part. Both are gated oncaching_applies(client, model), which requires the Orq router and an Anthropic (oragent/) model:cache_controlinside a content part is outside the direct OpenAI schema, so a client pointed atapi.openai.comor a self-hosted OpenAI-compatible endpoint is left untouched, and a routed non-Anthropic model gets no needless request-shape change (Orq documents the marker as ignored, not rejected, there). Anthropic caching is opt-in — without a breakpoint an append-only transcript is re-encoded in full every turn — while OpenAI, Gemini, DeepSeek and xAI cache automatically, so noprompt_cache_keyis set and no per-provider branch exists. On a routed client, marked content is sent as a text-block list rather than a plain string; a target that inspects the raw request body will see that shape. A cache write costs 1.25x, so neither helper marks an input whose text is belowCACHE_MIN_PROMPT_TOKENS(1024) — no Anthropic model caches below that, and the write would be pure loss. Covered: the simulation judge and user simulator, the red-teamOpenAIModelTarget, andOrqResponsesTarget— which is both the simulationagent:<key>target and the execution half of the default red-team agent backend. Not covered:redteam/runtime/jobs.pyandcommon/judge.py, which replay conversations throughexecute_chat_completionwithout breakpoints (RES-1360). Callers say how many trailing messages they rebuild each turn via the requiredvolatile_tailkeyword — marking a message that does not persist pays a 1.25x write nothing ever reads back, which is what the simulation judge's per-turn instruction did. Both paths are confirmed against live traces onanthropic/claude-sonnet-4-6with a uuid-salted cold prefix, three judgements each: Chat Completions read 0 / 6,991 / 7,881 tokens, Responses 0 / 7,417 / 8,304 (higher because its system prompt lives ininstructions, outside the marked input). - The adaptive red-team attacker loop is now cached too.
MultiTurnOrchestrator.run_attackmarks the adversarial system prompt and the end of the persisted transcript on every turn, gated oncaching_appliesexactly like the other surfaces. The call sits inside the turn loop deliberately:apply_cache_breakpointsreturns a copy and the transcript grows each turn, so a hoisted call would freeze the turn-1 snapshot and send it for the rest of the attack — a silent correctness bug, not a cost one, which is why the call-site test pins placement rather than just the marker count. Measured live onanthropic/claude-sonnet-4-6over five turns with a uuid-salted cold prefix: reads of 0 / 1,986 / 2,211 / 2,435 / 2,696 tokens against inputs of 1,988 / 2,214 / 2,438 / 2,699 / 2,948 — a 91.5% hit rate by turn 5, with only the newly appended pair paying full price. No public surface changes; an operator running the defaultopenai/gpt-5.6-lunaattacker sees no difference, since OpenAI caches on prefix without a marker. - A caller that rebuilds its trailing message every turn must say so.
apply_cache_breakpoints(messages, volatile_tail=N)keeps the breakpoint off the lastNmessages: mark one and the next turn puts transcript content at that position, the prefix diverges right after the system message, and the whole transcript pays a write that is never read. The simulation judge appends a per-turn instruction and passesvolatile_tail=1; so does the user simulator's first-message call. The Responses path takes the same guarantee throughmark_responses_input(input, volatile_items=N)— items, not messages, because oneMessagewith tool calls renders to severalinputitems;responses_volatile_items(messages, volatile_tail=N)converts between the two and is the only supported way to do so. - The simulation judge's system prompt is now byte-stable across a conversation. The
[ALREADY CONFIRMED]markers for settled criteria moved out of the criteria listing and into the trailing user message thatJudgeAgent.evaluateappends. The markers sat at token position 0, so everymark_settledcall invalidated the cached prefix — system prompt, whole transcript and tool schemas — to save a few dozen output tokens. The instruction is unchanged in substance and a custom judge is unaffected, but it now reaches the model later in the request; coverage is structural — no live-model test exercises the new placement. red_team()takesrecommendations=instead ofgenerate_recommendations=, and simulation gained the same flag with the same three forms:True(defaults),False(skip the LLM call), or a config instance —RedTeamRecommendationConfig/SimulationRecommendationConfig, both bounded andextra='forbid'. Two removals with no deprecation shim:generate_recommendations=onred_team(), and the long-deprecatedconfig=alias forllm_config=(the shim had targeted removal in 1.4.0 and outlived it).generate_focus_area_recommendations()likewise takesrecommendations=in place of itsmax_areas/max_tracespair. Rename the keyword at the call site; behaviour is unchanged. (RES-1286)- Simulation now generates remediation suggestions in-run rather than from a CLI post-run hook, so a saved run carries them regardless of caller.
simulate()/generate_and_simulate()default the flag toFalse, because the returnedSimulationResultlist has nowhere to carry suggestions — withsave/reportboth unset the run would pay for them and drop them, which now logs a warning.eq sim runandeq redteam runkeep--recommendationson by default. (RES-1286) - Assistant turns in a replayed transcript are sent to the Responses API as
output_textparts. A bare string — or a list ofinput_textparts — underrole: "assistant"is silently dropped by the Orq router (some backends 400 instead), so every stateless Responses target and the simulation judge/user-simulator were replaying history with the agent's own turns missing. The simulation judge saw a transcript with no agent replies and reported "the agent has not yet responded", which is the deeper reason no criterion about agent behaviour could ever fail (RES-1308). AffectsOrqResponsesTarget,OpenAIAgentTarget, red-team multi-turn replay, and simulation. An image part on an assistant turn is not representable and is now dropped with a warning. - The
criteria_metscorer returns 0.0 (was1.0) for a simulation that ended in an error or a timeout, and logs a warning. Such a run terminates before the judge audits anything, so its criteria outcome is unknown — scoring it a perfect 1.0 let a dead target inflate the run average andconversation_quality. A run with no criteria at all still scores 1.0. - Simulation
Scenariocriteria are now scored from an explicit per-criterion audit the judge returns on every turn (Judgment.criteria_verdicts), folded across the whole conversation. Previously pass/fail was inferred from the absence of a criterion id inrules_broken, somust_happencriteria could never fail andcriteria_metreturned1.0on every run (RES-1308). This changes scores for existing callers who use criteria:criteria_met,conversation_quality,rules_brokenandcriteria_resultscan now report failures where they previously reported none. Amust_happencriterion passes if it occurred in any turn; amust_not_happencriterion fails if it was violated in any turn. A customjudgethat does not emitcriteria_verdictsfalls back to the old behaviour, logs a warning naming the scenario, and is markedSimulationResult.criteria_verified = False.Judgment.criteria_verdictsislist[CriterionVerdict] | None.CriterionVerdict(new, public, inevaluatorq.simulation.types, re-exported fromevaluatorq.simulation) reportscriterion_id,occurredandevidencefor one criterion on one turn — occurrence only, never pass/fail.Nonemeans the judge reported nothing (unknown,criteria_verified=False);[]means it audited and had nothing left to report; a non-empty list is evidence. New public helperscriterion_id_for(index)andCRITERION_ID_PATTERN(both also re-exported fromevaluatorq.simulation) fix thecriteria_Nid format in one place. - The simulation judge's
finish_conversationtool no longer takes arules_brokenargument. Violations are derived in code from the occurrence audit andCriterion.type; the free-text list is the channel that could not fail amust_happencriterion in the first place, and asking for both gave them something to disagree about. A criterion the audit skipped now keeps its not-observed default instead of being rescued from free text.Judgment.rules_brokenandSimulationResult.rules_brokenare unchanged as outputs — only the tool input is gone, so a customjudgethat populates the field itself still works. - The judge stops re-auditing a criterion once it is confirmed to have occurred. Occurrence is sticky, so a settled criterion cannot change; it stays in the prompt (the judge needs it to decide whether to end the conversation early) but drops out of the per-turn
criteria_verdictspayload, which costs an id, a boolean and an evidence quote per criterion per turn. A custom judge without amark_settledmethod keeps auditing everything. metadata['criteria_meta']entries gainaudited— whether the judge actually returned an occurrence verdict for that criterion, as opposed to it falling to the not-observed default. Amust_happenthe judge confirmed never occurred and one it silently skipped both reportpassed: False; only this field separates them.Nonefor runs saved before the field existed.metadata['criteria_meta']entries also gainevidence— the quote from the turn where the criterion's occurrence first flipped, sourced from the judge'scriteria_verdictsaudit.''when the criterion never occurred,Nonewhen no tracker was available (same convention asaudited).- New
SimulationResult.criteria_verifiedfield, andcriteria_metreturns 0.0 for a run where it isFalse. It isFalsewhenever the judge returned no per-criterion occurrence audit for any turn — a customjudgepredatingcriteria_verdicts, or the built-inJudgeAgentterminating for safety after an unparseable tool call. Those verdicts came from the free-textrules_brokenlist, which cannot fail amust_happencriterion, so an all-green result there is unknown rather than passing; scoring it 1.0 reproduced RES-1308 one layer up, with a log line as the only signal.Noneon runs saved before the field existed, and those keep their previous score. - The
criteria_metevaluator now reportspass=False— with an explanation naming the cause — on exactly the runs it scores0.0: one that ended in an error or a timeout, and one withcriteria_verified = False. The flag was previously derived fromcriteria_metaalone, so an unaudited run landed on the evaluator trace span and the uploaded Orq experiment as a greenPASSbeside its own0.0. An errored run, which has nocriteria_metaat all, reported "No criteria defined for this scenario." andpass=Truefor a scenario that does have criteria. criteria_metno longer counts an individual criterion the judge never audited as met, on any surface. The scorer readsmetadata['criteria_meta']when present (the only placeauditedsurvives —criteria_resultsis keyed by description and carries no provenance) and counts a criterion only when it passed and was audited, logging a warning naming how many were not; the evaluator explanation printsUNKNOWN [required]: … (not audited)instead ofPASSfor it and excludes it frompass. This lowerscriteria_metfor runs where the judge audited some criteria and skipped others — previously the score counted the skipped ones as met while the report's own "N/M criteria met" tally did not, so the two contradicted each other. A criterion the judge settled early is audited (a verdict is what settles it), somark_settlednever costs a run a point;audited: None(a run saved before the field existed) still counts as met.auditedandevidencenow reach the reports.CriteriaRow(inevaluatorq.simulation.types, the per-criterion view model behind the report sections) gainsaudited,evidenceand a computedstateofpass/fail/unknown, andSimulationEntrygainscriteria_verified. Every surface rendersstate, notpassed: a criterion that passed only because the judge never audited it shows as not audited (a neutral?) in the dashboard, the HTML report and the markdown export, is excluded from the "N/M criteria met" tally, and a run withcriteria_verified = Falsesays so above the criteria list instead of showing a tally that contradicts itscriteria_metscore of0.0. The judge's evidence quote is shown beside the criterion it justifies.- A simulation whose target fails mid-run now keeps the criteria audit collected before the failure. The error result carries the folded
rules_broken,criteria_results,criteria_metaandcriteria_verifiedinstead ofrules_broken=[]and no metadata, so amust_not_happenviolation the judge confirmed on turn 2 survives the target dying on turn 4 — it previously vanished from the result and the report. (It never reachedfind_triggers: that helper returns[]for any errored result before it looks at criteria, and still does.) On this path only confirmed occurrence is knowledge: amust_not_happenthe judge saw violated stays failed, while amust_happenthat had not occurred yet is reported asunknown(row stateunknown,audited: False), never as failed — the run was cut short before that criterion had its chance, so folding the not-observed default would invent a failure the judge never made and add a phantom row to the cross-run failure-mode table. A target that dies before the judge audits anything reports every criterion that way, pluscriteria_verified=False. Such a run is still scored0.0bycriteria_met, because it terminated by error. EVALUATORQ_SPAN_MAX_TEXT_CHARSdefaults to capturing all message content (no truncation), in both the Python and TypeScript tracing layers. Set the env var to a positive integer (canonical:8192) to cap span text at that many characters (marker... [truncated]);-1,0, or unset all mean capture all. The cap applies uniformly to input and output message content. (RES-715 introduced an8192default; RES-899 reverts to capture-all and unifies the TS path, which previously hardcoded a separate2000-char cap.)evaluatorq()defaults todatapoint_parallelism=10(previously1). Evaluations are almost entirely provider-bound I/O, so the old default made the common case pay a latency penalty to protect the uncommon one. This changes behavior for existing callers who omit it: ten datapoints now run concurrently. Passdatapoint_parallelism=1to restore serial execution — do so if your provider rate-limits at low concurrency, or if your jobs mutate shared state that was previously serialized by accident rather than by design. Red teaming already defaulted to 10; simulation is raised from 5 to 10 in the same release, so the number now means the same thing on every entry point.- A target call makes exactly
max_target_retries + 1HTTP attempts, on every path.call_target_with_retryowns target retries, so the SDK's own budget is disarmed at that boundary (without_client_retries, awith_optionsclone — an injected client is never mutated and keeps its transport, auth, base URL, headers and timeout). Previously the two layers stacked and multiplied: an injected OpenAI or Responses client left at the SDK default made 9 HTTP calls where 3 were intended, at 3× the cost and latency, and a caller who setretry_counton such a client had no way to see it. Judge and pipeline calls are unchanged — each already had a single owner. Simulation agent calls (the user simulator and the judge agent) now honourLLMCallConfig.retry_count:SimulationAgent._call_chat_completions/_call_responsespassed no budget towith_retry, so they always used the module default of 5 transport attempts regardless of configuration. They now make exactlyretry_count + 1(default 2, previously 5). The chat-completions path additionally retries once within an attempt on an empty response — a content-level retry, so a model that keeps returning nothing still costs up to2 × (retry_count + 1)calls.retry_count/retry_on_codespassed for a target call are now ignored with a warning naming the owner rather than silently. evaluatorq()'sdatapoint_parallelismnow bounds evaluator fan-out too. Evaluators within a job previously ran with unbounded concurrency (a datapoint with 50 evaluators issued 50 concurrent provider calls no matter what the datapoint count said); they now share the same per-datapoint semaphore the jobs use. This lowers throughput for callers who relied on the unbounded behaviour — raisedatapoint_parallelismto restore it. The budget is shared, not split: a job releases its slot before its evaluators take theirs, so the two never contend anddatapoint_parallelism=1cannot deadlock.- New
llm_parallelism=onevaluatorq(),red_team(),simulate(),generate_and_simulate()andgenerate()— a ceiling on in-flight LLM requests for the whole run, counted per request rather than per task. Defaults to 10 when unset, and that default also covers LLM calls made outside any entry point (e.g. a standalonerun_pairwise()orsummarize_conversations()), so an unconfigured run cannot flood a provider. This lowers throughput for callers who relied on unbounded fan-out — pass a largerllm_parallelism(or-1for no ceiling at all) or wrap the call inevaluatorq.llm_concurrency_limit(n)to raise it. This is the knob to size against a provider concurrency limit:datapoint_parallelismbounds tasks, and the task bounds nest (datapoints × jobs/evaluators × jury width), sodatapoint_parallelism=10can mean anywhere from 10 to several hundred concurrent requests depending on the fan-out — the number was never something you could compute a request rate from. Requests routed throughcommon.llm_call(judges, juries, simulation agents, the red-team pipeline, the OpenAI backend) take a slot automatically; a job that calls a provider SDK directly is invisible unless you wrap it in the newevaluatorq.llm_slot()context manager, which is also what closes the gap for the ORQ and LangChain targets. Note this is a concurrency bound, not a rate limit: N slots isN / latencyrequests per second, so a provider that gets faster raises your request rate at a fixed N. - Nested LLM limits now stack. An inner
llm_concurrency_limit(20)inside a run capped at 10 stays under 10; an inner limit of 5 lowers the cap to 5.llm_concurrency_limit(-1)adds no cap and cannot lift an enclosing one. At the top level,-1still disables the default ceiling of 10. Unconfigured runs in the same event loop share the default budget, so their combined load stays under 10; explicitly configured runs have separate budgets. parallelism=is renameddatapoint_parallelism=onevaluatorq(),red_team(),simulate(),generate_and_simulate()andgenerate(), and--parallelismis renamed--datapoint-parallelismoneq redteamandeq simulate. Both old names still work and emit aDeprecationWarning;EvaluatorParamsaccepts either field name. With two concurrency knobs the bare name no longer said which one it meant — one counts datapoints, the other counts LLM requests. The OTel span attributesorq.redteam.parallelismandorq.simulation.parallelismkeep their keys, so existing trace queries and saved dashboard filters are unaffected. Breaking for hook implementors: the red-teamConfirmPayloadand the simulationSimulationRunMetakey is nowdatapoint_parallelism— a hook readingpayload['parallelism']willKeyError.- New
--llm-parallelismflag oneq redteam,eq simulateandeq simulate from-traces, exposing thellm_parallelism=ceiling to CLI callers.-1disables the ceiling. - First-message generation fails the datapoint instead of inventing an opening.
FirstMessageGenerator.generate()used to return a canned"Hi, I need help with: <goal>"when the model returned nothing usable or a 429/5xx outlasted retries, and that line was then simulated and judged as if it were real. It now retries an empty or refused reply up to 3 times, then raisesFirstMessageGenerationError(truncation raises at once, since the same budget would truncate again), and lets a provider error that outlasted retries propagate.DatapointGeneratorandsimulate()drop only the failed persona x scenario pair for unusable output or an exhausted transient provider error, log aGenerated N of M datapointswarning, and raise when every pair fails. Authentication, bad-request, and unexpected errors abort the batch so a configuration mistake cannot produce an incomplete dataset.datapoints_from_traces()still falls back to the trace's real recorded opening. Bothsimulate()and directDatapointGeneratorcalls batch first-message tasks at twice the active LLM ceiling, bounding pending tasks without adding another provider-request cap; a top-level-1leaves that fan-out unbounded. - Red-team reports show run warnings.
RedTeamReport.pipeline_warningsused to reach only the console. The HTML report, the Markdown report and the dashboard now show them in a "Run Warnings" section straight after the summary. A category whose strategy generation failed but which still ran its hardcoded strategies now gets a warning too, where before only a category left with zero strategies did. - Breaking: per-call-site concurrency knobs are removed, because the
llm_parallelismceiling now bounds those calls. Passing any of these raisesTypeError: max_concurrency=onrun_pairwise(),llm_jury_pairwise(),PairwiseComparatorandrun_jury().rate_limit_delay=andmax_concurrent_calls=onDatapointGenerator.generation_parallelism=onplan_strategies_for_categories()/plan_strategies_for_vulnerabilities(), anddatapoint_parallelism=ongenerate_dynamic_datapoints()/generate_dynamic_datapoints_for_vulnerabilities().
Delete the argument. To bound these calls, set llm_parallelism= on the entry point you run, or wrap a standalone call in async with evaluatorq.llm_concurrency_limit(n):. The paths these knobs capped (trace summarization and inference at 5, first-message generation at 5 plus a 100 ms delay, strategy generation at datapoint_parallelism) now share the default ceiling of 10. - Simulation's datapoint_parallelism defaults to 10 (previously 5), matching evaluatorq() and red teaming. This raises concurrency for callers who omit it; pass datapoint_parallelism=5 to keep the old value, or set llm_parallelism= to bound provider load directly. Applies to simulate(), generate_and_simulate(), SimulationConfig and eq simulate / eq simulate run. - evaluatorq() never exits the process when an evaluator reports pass_=False; it returns the results so library callers can inspect pass_ and choose their own gate. Red-team and simulation surfaces retain their own explicit failure gates. - loguru is now a core dependency (previously gated behind the [redteam] extra). This slightly widens the install footprint for non-redteam consumers but unifies the logging stack across the package. - openai (>=1.92.0) is now a core dependency (previously gated behind the [redteam] extra). The new llm_jury() evaluator imports it at package load, so every base install pulls it; this widens the base footprint for users who only call evaluate(), in exchange for llm_jury() working without an extra. - datapoints_from_traces() and extend_from_traces() now summarize every trace conversation unconditionally before the persona/scenario or traffic-profile call reads it — a short trace that previously skipped straight to that call now costs one extra LLM call in direct mode too. TraceAnalysisConfig.summarize_above_chars is removed with no deprecation shim; because TraceAnalysisConfig is extra='forbid', TraceAnalysisConfig(summarize_above_chars=...) now raises ValidationError — drop the field. A new summarize_conversations() entry point runs that summarize step directly: call it once and pass the result as summaries= to either function so a run that calls both does not summarize the same trace twice; a summaries= mapping is authoritative, so a trace_id absent from it (because it failed to summarize) is dropped rather than retried. (RES-1286)
-
W3C trace context propagation is now toggleable, and the Orq agent and deployment targets propagate it.
EVALUATORQ_PROPAGATE_TRACE_CONTEXT=false(or0) stops evaluatorq sendingtraceparent/tracestateon its outgoing LLM, Responses and target calls; the default stays on. Two target legs previously sent no trace context at all, so their server-side execution started its own root trace instead of nesting under the target-call span: the red-teamORQAgentTarget(agents.responses.create) and both deployment legs (deployment()/invoke(), used by the simulation adapter, and the red-team static deployment job). All three now pass the headers via the Orq SDK'shttp_headers. (RES-1422) -
llm_jury()andllm_jury_pairwise()now acceptpromptandcriteriatogether; the rule was exactly one of the two. It is now at least one: a panel with neither still raises. This exists because the two arguments no longer feed the same judge — an LLM judge reads the renderedprompt(which can pull the rubric in with a{{criteria}}placeholder), while a classify judge is handedcriteriadirectly as its question and never sees a prompt, so a mixed panel needs both. Nothing that was valid before is rejected now. Setting both on a panel that seats no classify judge and whose prompt never renders{{criteria}}logs oneWARNINGnaming that nothing reads the rubric, rather than raising.
Breaking Changes¶
- The deprecated Streamlit dashboards are removed:
eq redteam uiandeq sim uino longer exist. Both have printedWarning: eq … ui is deprecated; use eq dashboard instead.since July 2026;eq dashboardis the replacement and reads the same saved run files, with no migration beyond the command name. Theredteamandsimulationextras no longer pull instreamlit,plotlyorwatchdog— they now install onlyhuggingface-hubplus chart rendering (vl-convert-python) and, forsimulation,orq-ai-sdk. Theevaluatorq.redteam.ui,evaluatorq.simulation.uiandevaluatorq.common.uipackages are gone, includingsimulation.ui.token_display; nothing outside the deleted viewers imported them. red_team()parameter renamed:config=→llm_config=. The oldconfig=keyword was kept as a deprecated alias that emitted aDeprecationWarning; it has since been removed (see therecommendations=entry above).LLMConfigflat fields removed:attack_model,evaluator_model,adversarial_temperature,adversarial_max_tokens,llm_call_timeout_ms,llm_kwargs— replaced by role-basedattacker/evaluatorsub-configs (LLMCallConfig)wrap_simulation_agent()no longer accepts theevaluators=kwarg. Evaluators are wired throughevaluatorq()directly (the framework that consumes the job); callers passingevaluators=[...]will now get aTypeErrorand should move the list onto theirevaluatorq(..., evaluators=...)call instead (RES-594).simulate()andgenerate_and_simulate()no longer acceptagent_key=. The singletarget=parameter now selects the target:"agent:<key>"or a bare"<key>"(hosted Orq agent via the Responses router),"deployment:<key>"(legacy deployment), anAgentTarget, or a callable. Callers passingagent_key=...get aTypeError; migrate totarget="deployment:<key>"(ortarget="agent:<key>"). Theeq sim simulate/eq sim runCLI drops its matching--agent-keyflag — use--target deployment:<key>.simulate()andgenerate_and_simulate()now defaultupload_results=True. With the move to evaluatorq-native execution the framework's upload is the canonical persistence path — the previousFalsedefault left runs with no record anywhere. Setupload_results=Falseexplicitly to suppress (RES-594).AttackerResponseis removed — attacker LLM output is unified ontoevaluatorq.contracts.AgentResponse, the same shape targets and simulation agents already return.generated_promptis nowAgentResponse.text, andtruncatedis derivable fromfinish_reason == 'length'. Removed outright rather than left as a silent alias:AttackerResponse = AgentResponsewould have madeAttackerResponse(generated_prompt=...)quietly drop the prompt (AgentResponseignores unknown kwargs) and would have collapsedisinstancediscrimination between the two. ConstructAgentResponse(text=...)directly (RES-883).
Migration:
# Before
red_team(target, config=LLMConfig(attack_model="gpt-4o", evaluator_model="gpt-4o-mini"))
# After
from evaluatorq.redteam.contracts import LLMCallConfig, LLMConfig
red_team(
target,
llm_config=LLMConfig(
attacker=LLMCallConfig(model="gpt-4o"),
evaluator=LLMCallConfig(model="gpt-4o-mini"),
),
)
AgentTargetrelocated: moved fromevaluatorq.redteam.backends.basetoevaluatorq.contracts. Importing it from the old path now raisesImportError. TheBackendABC stays inevaluatorq.redteam.backends.base.AgentContext,ToolInfo,MemoryStoreInfo, andKnowledgeBaseInfoalso moved toevaluatorq.contracts, but — unlikeAgentTarget— their old import pathevaluatorq.redteam.contractsstill works (re-exported, same class objects,isinstanceunaffected). OnlyAgentTarget's old path is a hard break.
Migration:
# Before
from evaluatorq.redteam.backends.base import AgentTarget
# After
from evaluatorq.contracts import AgentTarget
AgentTargetunified onrespond(messages):respond(messages: list[Message]) -> AgentResponseis now the abstract method every target implements.send_prompt(prompt: str) -> AgentResponseis retained as a concrete back-compat shim on the ABC — it wraps the prompt in a single user message and callsrespond. Custom targets that previously implemented onlysend_promptmust implementrespondinstead.
Migration (bare custom subclass):
# Before — only send_prompt was abstract
from evaluatorq.contracts import AgentResponse, AgentTarget
class MyTarget(AgentTarget):
async def send_prompt(self, prompt: str) -> AgentResponse:
return AgentResponse(text=await my_llm_call(prompt))
def new(self) -> "MyTarget":
return MyTarget()
# After — respond is the abstract method; send_prompt is a free shim on the ABC
from evaluatorq.contracts import AgentResponse, AgentTarget, Message
class MyTarget(AgentTarget):
async def respond(self, messages: list[Message]) -> AgentResponse:
prompt = messages[-1].content or ""
return AgentResponse(text=await my_llm_call(prompt))
def new(self) -> "MyTarget":
return MyTarget()
OrqResponsesTarget is now stateless: __call__, _previous_response_id threading, _accumulated_usage, and get_usage() are removed. Conversation continuity is the caller's responsibility — pass the full transcript to respond each turn. Pass the target to simulate(target=...) (auto-routes to the target-agent path) or simulate(target_agent=...) instead of relying on __call__. Per-call token usage is reported on the returned AgentResponse.usage. - ORQAgentTarget last-user contract: respond(messages) forwards only the last user message to the ORQ agents endpoint (server-side state is held via task_id) and raises ValueError if messages[-1].role != "user". The endpoint, task_id threading, and usage accumulation are unchanged. - ChatMessage alias removed: the RES-596 deprecated alias ChatMessage = Message is gone. Import Message from evaluatorq.contracts (the public evaluatorq.simulation.ChatMessage re-export is also removed). - Simulation TargetAgent Protocol removed: the simulation runner consumes the canonical AgentTarget ABC from evaluatorq.contracts. The evaluatorq.simulation.TargetAgent / evaluatorq.simulation.runner.TargetAgent exports are replaced by AgentTarget. Migration:
# Before
from evaluatorq.simulation.types import ChatMessage
from evaluatorq.simulation import TargetAgent
# After
from evaluatorq.contracts import Message # ChatMessage was an alias of Message
from evaluatorq.contracts import AgentTarget # replaces the simulation TargetAgent Protocol
CallableTargetforwards the full transcript: the wrapped callable now receives the entire conversation as alist[Message](previously only the last user turn as astr), so stateless callables retain context across multi-turn attacks. The callable signature changes from(prompt: str)to(messages: list[Message]), andusage_fnfrom(prompt: str, response: str)to(messages: list[Message], response: str). The former last-turn-must-be-user guard is dropped (matching the other stateless targets). Callables that need OpenAI chat-completion dicts can callMessage.to_chat_completion()per element.
Migration:
from evaluatorq.contracts import Message
from evaluatorq.integrations.callable_integration import CallableTarget
# Before
target = CallableTarget(lambda prompt: my_agent(prompt))
# After — read the last turn off the transcript
target = CallableTarget(lambda messages: my_agent(messages[-1].content or ""))
New Features¶
eq redteam run,eq sim simulateandeq sim runtake structured input and give structured output.--config PATH(or-for stdin) carries the keyword arguments ofred_team(),simulate()orgenerate_and_simulate()as JSON, which reaches every data-shaped parameter, including ones with no flag such as inline personas and scenarios,attack_techniquesandedge_case_percentage. Unknown keys are rejected at every depth. A flag passed on the command line beats the file, field by field, and a field neither sets takes the SDK's default.--llm-config '<json>'merges into the LLM configuration, and the narrow model flags win over it for their own field.--jsonprints the finalRedTeamReportorSimulationRunon stdout and moves all other output to stderr, keeping exit codes.eq redteam schemaandeq sim schemaprint the JSON schema of the config file (--input, the default, with the CLI's own defaults) or of the result (--output). YAML is not accepted.- New
evaluatorq.signalspackage — computes deterministic ADR-25 structure, tool, autonomy and trajectory-tag signals from ATIF trajectories, with evidence, preconditions and approximation status; includes an evaluator adapter and opt-in Jev tool-role classification. Bundled tag thresholds are calibrated on local Claude Code sessions. - Added the trace finder surface.
eq findand the dashboard's Trace search page compile a natural-language question, select metadata filters with the classify model, search recent Orq traces through OQL, project each conversation, and classify matches with the classify model. Review-first mode, live progress, trace inspection, and JSON export are included. - Classify judges on a jury panel —
llm_jury()andllm_jury_pairwise()can now seat the Orq classify modeltypesafe/jev-latest(Jev) beside prompted LLM judges. A classify judge is not prompted: it is handed the material to judge and one question about it, and answers with a probability distribution — about $0.042 per million input tokens with output free, in under a second.criteriais its question and is required on a panel that seats one;labels={label: description}builds a choice question,levels=[...](2 to 10 ordered descriptions, numeric mode) builds a score question, and boolean mode compares the returned probability againstthreshold.state_fields=[...]picks which template paths are handed over as the material, defaulting to every placeholder the prompt renders exceptcriteria. The vote's explanation is synthesised from the numbers that decided it (noul=0.92 (threshold 0.5)), and the full distribution lands on the judge span asjudge.confidenceandjudge.probabilities. Settings a classify judge cannot use are named rather than dropped:system_prompt,prompt,temperature,structured_output,max_tokens,reasoning_effort,extra_kwargsandextra_bodyare listed in oneWARNINGwhen you set a non-default value, which says they still apply to prompted judges only when the panel actually seats one. Known classify ids warn at construction; models discovered through the fetched catalogue warn on their first call.repetitions > 1warns because a classify verdict is deterministic; on a pairwise panel this means repetitions of each ordering, not the two distinct swapped orderings. Routing is decided per call by the Orq model catalogue'ssupports_classifyflag, with a built-in list holdingtypesafe/jev-latestanswering (and warning once) only when the catalogue has no entry at all. The built-in list andregister_model()overrides provide synchronous hints for early validation and warnings, while the fetched catalogue remains authoritative at runtime, so a newly listed classify model works without manual registration. The endpoint lives on the Orq router, so a client that does not route through it cannot reach a classify judge. Red teaming and simulation are unchanged. llm_config=on the simulation entry points —simulate(),generate_and_simulate(),generate(),generate_personas()/generate_scenarios()(and their singular forms),summarize_conversations(),datapoints_from_traces(),extend_from_traces(),extend_from_experiment()andwrap_simulation_agent()now take a fullLLMCallConfig. It configures every simulation-side LLM call — five roles: the user simulator, the judge, the persona / scenario / first-message generators, the recommendations pass and the executive summary — and never the target under test, which is configured where it is constructed.max_tokensis narrower than the rest: the user simulator, the judge and the executive summary read it, while everygenerate_structuredcall site — the generators, the recommendations pass and the trace helpers — sizes its own budget from the item count and logs that the config's value was not used.llm_config.clientis honoured on every one of those paths, so a caller who supplies their own client needs no credentials in the environment. Not exposed on the CLI:eq simulatestill takes only--sim-model, which builds anLLMCallConfignaming that model, and the rest of the config is a Python-API surface. This is the seam that was missing onceLLMCallConfig.temperaturelost its default: a caller who wants to pin sampling on the judge (or setreasoning_effort,timeout_ms,extra_body) now has one object to pass instead of hand-building each agent.llm_configis the only place the simulation-side model is named — thesim_model=keyword that once shared that job is removed, see the entry under Notable defaults. Only the fields you set take effect, so an unsettemperaturestill means the parameter is omitted from the request entirely. Mirrorsred_team(llm_config=LLMConfig(attacker=..., evaluator=...)), which already had the equivalent surface. The user simulator and the judge take that config plus their prompt data as separate arguments, so a value you set explicitly toNone—temperature=None, meaning "send no temperature" — reaches the request instead of collapsing back into "unset"; the deprecatedAgentConfigand its subclasses are still accepted and still cannot express that distinction.- Token-usage tables in the red-team and simulation reports now report cached input tokens as a share of input (
75 (75% of input tokens)) rather than a bare count. A raw cached figure says nothing about whether caching worked without the denominator it was cached against. Red-team reports gained the row outright — they carried only cache writes, which the Orq router reports as permanently 0, so the one measurable cache number was missing from the surface an operator reads. The row is omitted entirely when nothing was cached, so0 (0% of input tokens)never reads as a measured result. This is a token share, not a cost share — a cached token bills at roughly a tenth of a fresh one and the provider exposes no cached-vs-uncached cost split, so the row says "of input tokens" and sits in the token table rather than besideTotal Cost. Simulation's row was relabelled fromCached Tokens (retrieved)toCached Input Tokensso both surfaces name the same number the same way. extra_body=on the jury helpers —llm_jury(),llm_jury_pairwise()andPairwiseComparatornow takeextra_body=alongsideextra_kwargs=. These helpers build theirLLMCallConfiginternally, so body fields previously had no seam on the jury path at all:extra_bodyis a structural field the call site owns, which means routing it throughextra_kwargsraises — and since a judge failure becomes a verdict rather than propagating, the symptom was a judge that always failed.extra_kwargsstill replaces a top-level call argument;extra_bodyis merged per key, so router-owned body fields survive alongside yours.llm_jury()— LLM-as-a-jury evaluator forevaluatorq(evaluators=[...]). A single judge or a panel rates a target output against criteria; verdicts can be boolean (default), labeled categorical (labels=+passing_labels=), or numeric (verdict_kind="numeric"+threshold=). The panel consensus rule is selectable viaaggregator=:"mode"(default) or"majority"(strict >50%) for categorical,"mean_std"(default) /"median"/"min"/"max"for numeric, or a customCallable[[list[JuryVote]], ...]. Uses structured generation (tiered.parse→json_objectfallback) and resolves the LLM client lazily on first scorer call so declaring an evaluator never requires credentials. The Responses-API path is deferred (RES-972). (RES-848)OWASP_LLM_TOP_10andOWASP_ASI_TOP_10— publiclist[str]constants exported fromevaluatorq.redteam. Pass them tored_team(categories=OWASP_LLM_TOP_10)to run a full framework sweep without spelling out individual category codes (RES-815).simulate()andgenerate_and_simulate()accept a new opt-inupload_results=flag (defaultFalse). When set toTrue, results are uploaded to the Orq platform after the run, surfacing as an experiment whenORQ_API_KEYis configured. Upload errors are logged but never fail the call. Both functions also acceptevaluation_description=andpath=parameters mirroringevaluatorq()(RES-598).LLMCallConfig— per-role LLM configuration withmodel,temperature,max_tokens,timeout_ms,extra_kwargs, andclientfieldsLLMConfig— now role-based viaattacker: LLMCallConfigandevaluator: LLMCallConfig; retry, cleanup, and target-agent timeout settings retained at top levelLLMCallConfigexported from theevaluatorq.redteampublic APIOpenAIModelTarget.send_promptnow enforcestimeout_msviaasyncio.wait_for- Evaluator role config (
temperature,max_tokens,timeout_ms,extra_kwargs,client) fully propagated throughOWASPEvaluator,create_dynamic_evaluator, andcreate_owasp_evaluator simulate()andgenerate_and_simulate()accept newevaluation_description=andpath=parameters, forwarded straight toevaluatorq()(RES-598).simulate()andgenerate_and_simulate()now run on top ofevaluatorq(): persona × scenario datapoints are materialised, executed via a single evaluatorq job, and scored via adapted evaluators. This brings auto-upload, OTel tracing, the results table, CI gating, and dataset-id support to the simulation entry points "for free". The bespoke parallelism loop was removed;simulation/upload.pyis kept as a standalone helper for direct callers but is no longer invoked fromsimulate()(RES-594).simulate()accepts a newdataset_id=parameter — when set, simulation datapoints are streamed from the named Orq dataset (each row'sinputsmust already match a simulation input shape) instead of being passed inline. Mutually exclusive withdatapointsandpersonas/scenarios(RES-594).simulate()andgenerate_and_simulate()accept a newexit_on_failure=parameter, defaultTrue, for their own dropped-row gate. Evaluator score failures are returned in the results; dropped jobs raiseSimulationDroppedError. Passexit_on_failure=Falsefor interactive / exploratory runs where you want dropped rows surfaced as warnings + error metadata instead of a non-zero exit (RES-594).
Bug Fixes¶
- A call is now priced against the model that answered it rather than the id the caller asked for, falling back to the requested id when the served one is not in the catalogue, except for an
orq/*router, whose call stays unpriced with a warning instead of billed at the router's rate. Anorq/*router resolves to a different model per request, and the catalogue lists a single nominal rate for the router itself, so a run through one was billed at that headline rate no matter what served it. Measured on the live catalogue:orq/autorouter-anthropic-balancedresolved toeu.anthropic.claude-haiku-4-5and was costed 4.55x over. When the served model cannot be priced, a router call now logs a warning that says why instead of reporting no cost silently. A prompt-cache breakpoint is now placed for anorq/*model too, for the same reason it is foragent/<key>: the resolved model is not knowable when the marker is written, and the router ignores rather than rejects the marker (RES-1573). safe_substitute()dict keys were broken by Ruff RUF027 auto-fix inattack_generator,capability_classifier, andobjective_generator— LLM prompts were receiving unsubstituted{placeholder}text, silently producing degraded attacksgenerate_recommendations=Truenow correctly usesllm_config.evaluator.clientbefore falling back tocreate_async_llm_client()- All hardcoded timeout literals (
240_000,90_000) replaced with config-driven values fromLLMConfig/DEFAULT_TARGET_TIMEOUT_MS OpenAITargetFactorynow propagatesmax_tokensandtimeout_msto created targetsLangGraphTargetnever populatedToolCallOutputItem.result: LangGraph returns tool results as separateToolMessages andrespond()skipped every message that was not anAIMessage.render_tool_calldrops a call whose result isNone, so judges saw transcripts with zero tool calls and any "did the agent use tool X" criterion failed regardless of what the agent did. Results are now paired with their call within the turn, as the pydantic-ai and OpenAI Agents targets already did. Two shapes still drop, both now warning: a call whoseToolMessagearrives in a laterainvoke(interrupt/resume graphs), and a call the graph emits without anid. AToolMessagecarrying Anthropic-style content blocks is unwrapped to its text rather than rendered as the block envelope, and a message dict keyed ontype(BaseMessage.model_dump()) is now read rather than silently skipped.- A tool-call id from a non-OpenAI provider (
toolu_*on Anthropic via LangGraph,run-*via pydantic-ai) was replayed as a Responses-APIfunction_call.id, which rejects the whole request with "Expected an ID that begins with 'fc'" — killing the simulated user and the judge mid-run on any OpenAI-family model. Foreign ids are now dropped by the newresponses_function_call_item_id()helper (exported fromevaluatorq.openresponses.input_items).call_idis untouched and still carries the provider's id: it, notid, is what pairs a call with its output, and the API assigns an item id when none is sent. new()on every target (LangGraphTarget,CallableTarget,CrewAITarget,VercelAISdkTarget,PydanticAITarget,OpenAIAgentTarget,OrqResponsesTarget,OpenAIModelTarget,ORQAgentTarget) constructs viatype(self)instead of a hardcoded class name, so a subclass no longer silently degrades to the base class on every parallel job.
Internal¶
create_model_jobsplit down tocreate_deployment_job: the router-model leg it also built was unreachable.parse_targetreturns onlyAGENTorDEPLOYMENT—llm:/openai:/direct:and unknown prefixes all raise — so_create_job_for_target's trailingcreate_model_job(model=value)could not execute, and thereasoning_effortforwarding inside that leg was dead with it. The fallback now raises for a kind with no leg rather than routing it to a job no target string could produce. Not exported from any__init__, so this is internal surface only.usage_from_exceptionincommon/structured_output.pyreplaces five hand-copiedgetattr(exc, 'usage', None)sites, each of which carried its own copy of the harvest rationale. A guardrail intests/test_reuse_guardrails.pynow fails on a sixth.SaveModeconverted fromLiteraltoStrEnum- Timeout defaults centralised in
contracts.py(DEFAULT_TARGET_TIMEOUT_MS = 240_000);PIPELINE_CONFIGimport removed fromopenai.pyandregistry.py MultiTurnOrchestrator.llm_kwargsconstructor param deprecated — merged into_cfg.attacker.extra_kwargsat init time; useLLMCallConfig.extra_kwargsinstead- RUF027 added to Ruff ignore list (intentional literal string keys used as
safe_substitutetemplate placeholders) - CLI
--saveflag migrated totyper.Choice - Ruff cleanup across all redteam modules (import sorting,
Optional[X]→X | None,TYPE_CHECKINGguards)
Breaking Changes (RES-877)¶
AgentTarget.send_promptremoved:respond(messages: list[Message]) -> AgentResponseis now the sole response method on every target; callers own the conversation transcript. Migratetarget.send_prompt("x")totarget.respond([Message(role="user", content="x")]).OpenAIModelTarget,VercelAISdkTarget, andOpenAIAgentTargetare now stateless: per-instance_historyis gone. Multi-turn conversation state is owned by the red-team orchestrator, not the target.evaluatorq.redteam.ErrorInforenamed toRunError: update any imports orisinstancechecks that reference the old name.
Migration:
# Before
response = await target.send_prompt("Hello")
# After
from evaluatorq.contracts import Message
response = await target.respond([Message(role="user", content="Hello")])
New Features (RES-877)¶
AgentResponseError— a per-response error marker exposed onAgentResponse.error; used by the orchestrator to exclude failed turns from the replayed transcript.turns_to_messages(turns, *, skip_errors=False)— helper exported fromevaluatorq.redteam.contractsthat converts a list of completed turns into a flatlist[Message], optionally dropping turns whose response carries anAgentResponseError.classify_error_type(error, *, existing_type=None)— exported fromevaluatorq.redteam.contracts; infers a coarseerror_type(content_filter,rate_limit,timeout,network_error,server_error,client_error, orunknown) from an error string. Shared by the orchestrator and report converters. On a per-responseAgentResponseError, the orchestrator records an unmatched (unknown) result astarget_error, so that field never carriesunknown.- Tool-call fidelity on replay — the transcript replayed to a target now preserves assistant
tool_callsandtoolresults across turns (OpenAIModelTargetas OpenAI chat params,VercelAISdkTargetas AI SDK CoreMessagetool-call/tool-resultparts,OpenAIAgentTargetas Responses-APIfunction_call/function_call_outputitems), so multi-turn tool-using agents see their prior tool context.VercelAISdkTargetacceptsmessage_format="v5"(default) or"v4"to match the endpoint's AI SDK version (input/output:{type,value}vsargs/result). Errored turns recorded by the orchestrator now carry a classifiedAgentResponseError.error_typeinstead of a flattarget_error.
Internal (RES-899)¶
- Unified tracing layer: the generic OTel span-recording helpers previously duplicated across
redteam/tracing.pyandsimulation/tracing.pynow live in a singleevaluatorq.common.tracingmodule (truncate_for_span,capture_message_content,record_token_usage,record_llm_response,record_llm_input/output,set_span_attrs,get_trace_context_headers). Domain-specific span builders (with_redteam_span,with_simulation_span,with_llm_span) stay in their domain modules and import the shared helpers. The common module never imports fromredteam,simulation, oropenresponses.
Changed (RES-899)¶
- Span PII gate env var renamed to
EVALUATORQ_CAPTURE_MESSAGE_CONTENT(defaulttrue), replacing the previousOTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT. The same name now gates both the Python and TypeScript simulation/red-team tracing layers. Setfalse/0to keep raw prompt and response text off spans (token usage, model, finish reason, and latency are still recorded). - Span text truncation defaults to capture-all in both Python and TypeScript.
EVALUATORQ_SPAN_MAX_TEXT_CHARSis unset by default (no truncation); set a positive integer (canonical:8192) to cap input and output message content, with the shared... [truncated]marker.-1/0/ unset all mean capture all. The TypeScript path previously hardcoded a separate2000-char cap with a…marker — both are gone.
Fixed (RES-899)¶
retry_statusesaugments the default set again: passing a custom set (e.g.{429}) no longer silently drops the built-in429 + 5xxretries — the custom statuses are added to the defaults, not substituted for them. (This restores the intended RES-897 review behavior, which was lost when #150 merged without the fix.)