Skip to content

Changelog

All notable changes to evaluatorq are documented here.


[Unreleased]

  • Four model roles, fast, smart, classifier and embedding, name which model does which kind of work, with global flags to set them. eq takes --fast-model, --smart-model, --classifier-model, --embedding-model and a repeatable --model-override TASK=MODEL ahead of the subcommand; EVALUATORQ_FAST_MODEL, EVALUATORQ_SMART_MODEL and EVALUATORQ_EMBEDDING_MODEL join EVALUATORQ_CLASSIFIER_MODEL; the settings file gains fast_model, smart_model, classifier_model, embedding_model and model_overrides; and dashboard Settings shows one field per role. An unknown --model-override task exits with the list of valid tasks. A settings file saved before this change still loads: a non-default compiler_model or apply_model becomes a task override, and a value equal to the old default is dropped. The order of precedence and the task list are in Configuration › Models.
  • Deprecated: EVALUATORQ_COMPILER_MODEL and EVALUATORQ_APPLY_MODEL. Both still pin their task (finder.compiler and apply) and log one warning per process. Use --model-override, model_overrides in the settings file, or the role variables.
  • The Insights run form lets you choose the summary, classifier, embedding and (for the Question source) question compiler models from the Orq catalogue. A run saves the compiler model in its config so Re-run prefills it, and starting a run is rejected when the catalogue says the classifier cannot serve /classify or the embedding model is not an embedding model; with no catalogue the typed value is accepted and a warning is logged.
  • The Insights new-run page, the run-page "+ New run" dialog and "Re-run" now show one server-rendered three-step form (Traces, Analysis, Review). It combines both earlier forms: Presets, every source and label either form offered, individual coding-agent question toggles, custom questions, and Browse uploads for Finder exports and snapshots. A rejected submission re-renders with the entered values and the error, and review pages for saved runs now show the sign-in warning.
  • The form estimates traces, cost and time before the run starts. The trace count is exact for a Finder export or snapshot and an up to ceiling for Recent and Question, taken from the Orq per-value counts and the trace limit. Cost multiplies traces by capped tokens per stage by the stage model's catalogue price; time uses the median per-trace seconds of your earlier Insights runs. Anything without a basis reads unknown with the reason, and a total that leaves stages out says priced stages only or timed stages only.
  • Filter menus on Traces, Trace search and the Insights form show a trace count beside each value. Counts are Orq's for the search window, or a tally of the loaded rows on Traces, and values sort from most to least frequent. All three pages render the menu and chips through one facet_picker module.
  • Built-in Top 10% views on /traces for slowest, costliest, token-heavy and largest-context traces and for costliest, token-heavy and longest conversations (BOPS-1243).
  • New search on Traces reloads the table with its searched traces, so AI matches are visible in the table and related views.
  • Ask AI can ask up to three questions of each trace, or none. The compiler now produces zero to three named classifier dimensions. A trace is included only when it matches all of them, and each dimension gets its own column in the Traces table, the Trace search table, and eq find output. All of a trace's dimensions share one classify call, and a question that filters alone answer (such as Any traces above 50k tokens) runs no per-trace classifier at all. Within results now applies the question's metadata filters and token or duration bounds to the loaded rows. The run export moves to schema_version: 2: a dimensions list replaces the top-level task and selection, and each trace carries per-dimension answers in place of value, confidence and probabilities.
  • Ask AI progress is now visible while traces load and classify. The /traces progress line shows live load counts, a thin progress bar appears for loading and classification, and the corner status badge is hidden while a run is active.
  • Within results questions set the table's filters. A bound or facet in the question (such as over 20k tokens) narrows the Traces table to the matching loaded rows and shows up as a filter chip, in the Filters count and in the Filters menu; when no loaded row qualifies the run completes with zero rows and a notice instead of failing. Judging is capped by the AI trace limit in settings rather than by the number of rows in the table, AI answers stay on their rows when the table reloads until Clear AI results, and a wide table scrolls sideways inside the card with a pinned header and time column.
  • Ask AI only counts the answers a question asks for. The compiler no longer selects every choice label (for example Neither) or the healthy label for a negated question; such a plan is rejected and retried once, so the match count reflects the traces the question is looking for.
  • Trace finder planning and filtering now explain more of what happened. Descriptive phrases such as coding agents become classifier dimensions instead of being treated as numeric-only questions, boolean selections accept true and false labels, model or provider filters can find model-level traces, and Within results names the filter and nearest loaded value when every row is dropped. Numeric-only plans with uncovered words show a warning that names the uncovered question text.

Notable defaults

  • Red-team attack generation, all judges and evaluators, and apply-recommendations now default to the smart role, openai/gpt-6-sol (was openai/gpt-5.6-luna). This covers red-team attackers and evaluators, the simulation judge, the default llm_jury() panel, the dashboard's apply merge and Insights summaries. It raises the cost of red-team, simulation-judge and apply runs. Set EVALUATORQ_SMART_MODEL=openai/gpt-5.6-luna to keep the old model. The simulated user, simulation generators and the trace-finder compiler use the fast role, openai/gpt-6-luna (was openai/gpt-5.6-luna), and DEFAULT_PIPELINE_MODEL exports the fast role default. A call that passes model= explicitly is unchanged. With only OPENAI_API_KEY set, set EVALUATORQ_FAST_MODEL=gpt-6-luna and EVALUATORQ_SMART_MODEL=gpt-6-sol, since the built-in ids are provider-prefixed.
  • simulate() and generate_and_simulate() now generate remediation recommendations by default (recommendations=True), matching red_team() and the CLI. That is one extra LLM pass per run over the results with a fixable failure. The returned SimulationResult list cannot carry them, so a call without save=True or report_path= generates them, discards them and logs a warning; pass recommendations=False to skip the call.
  • The minimum supported pydantic is now 2.12. --config and --llm-config validation relies on per-call extra='forbid' (model_validate(..., extra='forbid')), which first shipped in 2.12, to reject misspelt keys at every depth, including inside models that ignore extra keys by default.
  • eq sim simulate and eq sim run leave --name, --max-turns and --datapoint-parallelism unset unless passed, so the SDK's defaults apply. The effective turn cap and concurrency are unchanged (10 and 10). Without --name, an uploaded experiment takes the SDK's generated simulation-<timestamp>-<id> name instead of sim; the saved run is still named sim. The CLI now saves through the SDK's save=, so a saved CLI run also gets a run manifest, and --save/--no-save replaces the lone --no-save (which still works).
  • eq redteam run gains --generate-strategies/--no-generate-strategies and --cleanup-memory/--no-cleanup-memory pairs; the existing --no-* spellings still work. Every flag that has a --config field now takes its default from the SDK, through one shared constant where the SDK resolves None (DEFAULT_DATAPOINT_PARALLELISM).
  • The Insights run form no longer has typed server-path inputs for Finder exports and snapshots; use Browse. The dashboard accepts Finder exports up to 10 MiB and snapshots up to 100 MiB; larger files go through eq insights --from-finder PATH or --from-snapshot PATH. The single "Analyze coding agents" checkbox is replaced by per-question toggles, and a dashboard run now uses the models chosen in the form rather than only the defaults.
  • Balanced Trio and Single-Provider Trio replace GPT-5.6 Luna with openai/gpt-6-luna at medium. The successor is both stronger on the same Intelligence Index v4.3.2 and cheaper at captured $0.10/$0.50 input/output rates per million tokens. Their flat estimated costs fall to $16.36 and $57.90 per 1,000 items respectively; the EU-hosted Luna seat is unchanged.
  • Open-Weight / Portable now seats wafer/Kimi-K3 at the host's default effort, none, instead of Baseten at max. This accepted host-and-effort trade lowers the flat estimated panel cost from $37.66 to $34.28 per 1,000 items, using Wafer's captured $3.00/$12.75 input/output rates per million tokens. It does not establish equal benchmark accuracy at the two reasoning settings.
  • The Traces page (/traces) automatically loads up to 200 traces from the last seven days with no user-entered filters and no AI calls; a saved project selection still applies. Press Load to fetch again with the selected controls. The table supports up to 5000 rows, and its default columns are Status, Started, Agent, Model, Tokens in, Tokens out, Cache read %, Cost, and Duration; the selected columns are saved in the dashboard settings file (default .evaluatorq/dashboard-settings.json, configurable with EVALUATORQ_DASHBOARD_SETTINGS) under explorer_columns.
  • Ask AI on Traces defaults to "Just proceed" and, when traces are already loaded, to the "Within results" scope, which classifies the loaded population instead of starting a fresh search. Choose Review first under Ask AI on traces in Settings to see the plan before any classifier call is billed; asking within more than 500 loaded traces is forced through review regardless of that setting.
  • Trace search remains available as a separate dashboard page at /find; the new Traces explorer is at /traces. Both pages share the facet menu, filter chips, and trace drawer while showing tables suited to their result types.
  • ModelInfo costs can now be None. The catalogue used to drop any model Orq lists without a usable USD price; it now keeps the entry with input_cost_per_1k and output_cost_per_1k set to None, so the dashboard Settings model picker shows it. Calls to such a model stay unpriced, as before. Code that does arithmetic on the result of get_model_info() must check for None. ModelInfo also gains a trailing model_type field.
  • New evaluatorq.formats package — ChatConversation, ResponsesConversation, OtelTrace and AtifTrajectory convert between any two formats via ATIF; ATIF is read at v1.7 and v1.8 (older versions raise; AtifTrajectory.from_json also rejects a document with no schema_version, while a trajectory built in code defaults to v1.7) and written with the trajectory's own version (v1.7 for converter output), upgraded to v1.8 when a trajectory holds audio.
  • Jury presets run at each judge's default reasoning effort, and published costs assume output tokens equal to input. JuryPreset.seated_efforts() now reports the catalog default for each seat (mostly medium), which is the effort a preset call actually runs at, since presets send none; seats are chosen in this package at that effort rather than at each model's top effort. Published estimated_cost_per_1k now assumes a flat 1,500 input and 1,500 output tokens per seat rather than 1,500 and 150, so every figure is about four times higher: Balanced Trio 17.56, Strong Jury 85.50, Open-Weight / Portable 37.66, EU Region 27.75, Single-Provider Trio 59.10. JuryPreset.priced_below_seated_effort() is removed, and the rate table's seated_effort and priced_at_ceiling fields are replaced by default_reasoning_effort. Strong Jury seats anthropic/claude-opus-5-5 in place of anthropic/claude-opus-5, and Balanced Trio's reserve is deepseek/deepseek-flash, the id that now serves DeepSeek V4.1 Flash. ModelInfo gains default_reasoning_effort.
  • New evaluatorq.backends.CodingAgentTarget runs Claude Code, Codex CLI or OpenCode as the system under test, directly or through orq launch, on the host or in a reusable per-clone container. Coding-agent turns now have a 300 s idle limit under a 2 h hard cap in host mode too; a host-mode agent that keeps printing can run for up to 2 h, where the previous 210 s effective ceiling cut it off. Container runs default to each agent's bypass mode and block Linux privilege escalation. Failures surface as cli.* error codes, including non-retryable timeout and container startup errors. SimulationRunner passes the target's own map_error to the retry helper, so backend-specific codes reach simulation results where the default mapping used to apply.
  • A static red-team row with no user message now fails instead of being sent. The plain static pipeline used to send an empty prompt to an AgentTarget or agent: target and score its reply, so a malformed dataset row could come back RESISTANT. It now raises before any target is created and appears under Failed Jobs, which the hybrid pipeline's static leg already did.
  • CodingAgentTarget now runs on a Windows host. Container mode used os.getuid() and SIGHUP, the orphan sweep probed liveness with os.kill(pid, 0) (a Ctrl+C on Windows), and killing a timed-out turn used os.killpg(); the first container turn raised AttributeError, a leftover container made every start raise OSError, and a host-mode timeout left the agent running. On Windows the container now runs as the image's agent user (uid 1001), the sweep asks the kernel whether the owner is alive, and a timed-out turn ends its process tree with taskkill /T /F, best effort. Container cleanup there relies on atexit and the lease, since SIGTERM and SIGHUP are not delivered. The agent binary is now resolved on the PATH from env, so an npm-installed claude.cmd is found on Windows; a .cmd or .bat launcher is refused with cli.unsafe_shim when an argument carries characters cmd.exe would interpret.
  • Insights labels and summaries now read a compact view of the conversation instead of the trace projection. The old projection kept the end of long traces within its byte budget and dropped the first user message in 82 of 100 sampled coding-agent traces; the new view keeps every opening user message, the start and end of each assistant message, and one line per group of tool calls, and leaves tool outputs out. Labels and summaries on long traces can change.
  • The built-in user_frustration label now rates 1 (calm) to 5 (angry or giving up) instead of six levels from 0, so saved scores are not comparable with earlier runs.
  • Insights summaries no longer produce tools_used; each trace instead records tool, shell-program, and skill counts taken from its messages as tool_stats.
  • eq insights --coding and insights(coding_analysis=True) add coding-agent labels (task type, outcome, verification, scope creep, user corrections, unfixed errors, risky actions) to traces the classifier identifies as coding agents. Off by default.
  • Insights --classifier-model now follows EVALUATORQ_CLASSIFIER_MODEL or the dashboard setting when either is configured; typesafe/jev-latest remains the fallback.
  • Trace projections used by finder classification include each paired tool result's completion status and a fixed diagnostic category for recognized failures. Raw tool result bodies are omitted because they may contain credentials that pattern-based redaction cannot reliably identify, so classifications on tool-heavy traces can change. Insights allows more output tokens for summaries and cluster names after a live tool-heavy run exhausted the previous limits.
  • eq find --positive-only keeps only matching trace records in its JSON export. The terminal table still shows matches and a full-run summary, JSON counts describe the full run, and debug diagnostics still show every classifier call.
  • Trace Insights adds insights(), eq insights, and a read-only dashboard page for analyzing saved Orq trace populations. Runs keep filtered population selection separate from fixed classifier labels and clusters discovered from summary text. It defaults to openai/gpt-6-luna for summaries and stores per-item summaries and embeddings in .evaluatorq/cache/insights.sqlite; install evaluatorq[insights] for the numerical dependencies and use cache=False or --no-cache to bypass the local cache. The concerning preset uses classifier level indices 0–4. In the CLI, the priority matrix uses intent by default when that dimension is selected, or the first selected dimension otherwise. The dashboard accepts Finder exports up to 10 MiB when starting an Insights run.
  • Saved Insights runs now open in a four-view review page by default. Themes, Activity, Map, Compare, and the trace list share filters that the URL restores. Activity reports recorded tool, shell-command, and skill counts, and distinguishes traces with missing activity data from traces with no recorded use. The dashboard Map is a display-only 2D projection of saved coordinates; it does not rerun embeddings or change the run. The new-run sheet supports validated Finder and snapshot uploads up to 100 MiB, individual coding-label selections behind the coding_agent gate, and validated custom questions. Existing saved runs, exports, trace details, and legacy tab URLs remain available.
  • Insights runs save recorded token usage and known provider cost for tracked label, summary, describe, merge, and embed stages; semantic-query population compilation and filter-selection calls are excluded, and failed calls may have unknown billing. Missing or unpriced usage and excluded calls mean displayed dollars can be below whole-run spend. The dashboard shows known cost and partial price coverage in the run header, plus live trace counts for label and summary progress; a partial dollar amount is a lower bound.
  • eq dashboard now reads .env from the launch directory and opens the local dashboard in a browser when the server starts. Already exported environment variables take precedence; the browser opens once, not on each hot reload.
  • Dashboard Settings can save an Orq profile, workspace, and project together. The selected profile supplies the API key and host, its workspace slug supplies Orq trace links, and the selected project ID limits dashboard Trace search. Changing profiles refreshes the available workspace and projects before saving. The CLI's masked profile keys are resolved from its private local credential file; no key is written to dashboard settings. The eq find CLI uses the saved profile and project by default; --profile and --project override them for one run.
  • The trace finder defaults to a seven-day window, a 500-trace population default (maximum 5000), 100-way classify concurrency, openai/gpt-5.6-luna for compilation, and typesafe/jev-latest for classification. Override the models through the dashboard Settings page; set the window, limit and parallelism per run on Trace search, through the finder flags, or through the documented environment variables.
  • eq find now shows a terminal activity indicator without printing every poll. Pass --debug to print changed progress and the compiler and classifier requests and responses, including trace content.
  • inference no longer has to be passed with a recorded-output source. It now defaults to unset and resolves from data: ExperimentInput and TraceInput are replay sources by construction and resolve to inference=False, anything else to True. evaluatorq('run', data=TraceInput(limit=50), evaluators=[...]) is now the idiomatic call; it used to raise, and the only way to learn the required flag was to trip that error. An explicit inference=True with a replay source still raises, and inference=False everywhere it already appears keeps working unchanged.
  • TraceInput.limit is now int | None, defaulting to None, and cannot be combined with trace_id=. Trace mode never read limit, so TraceInput(trace_id='t1', limit=50) used to validate and then return exactly one trace with no warning; it is now rejected. The resolved count (still 20 when unset) is the query_limit property — read that rather than the field.
  • A trace query that selects nothing is no longer a green run. fetch_traces warns when a query matches no traces, evaluatorq(data=TraceInput(...)) raises rather than evaluating zero rows, and red_team(datapoints=[]) is rejected instead of running zero attacks and exiting 0. A mistyped search= used to produce a successful-looking report with a resistance rate computed over nothing Simulation matches them: simulate()/datapoints_from_traces() raise when a TraceInput yields no conversation with a user turn, where they used to return an empty list and report a run of zero personas.
  • One unreadable trace no longer kills the batch. redteam.datapoints_from_traces skips a trace that failed to import (or has no seedable turn) with a warning and keeps the rest, matching what simulation already did and what fetch_traces always documented; it raises only when no usable trace remains. Failed imports still become individually-failing rows on the core evaluatorq() path. All three surfaces now partition through the shared common.trace_input.partition_traces.
  • The trace import now retries rate limits and server errors. The Orq trace endpoints have no SDK method, so fetch_traces calls them over raw httpx; with_retry recognised only the OpenAI SDK's own error classes, so a 429 or a 503 raised by raise_for_status() ended the import on the first attempt while the docstring promised retry. httpx.HTTPStatusError is now retried on the same status rule as every other call path.
  • A naive TraceInput.start_time / end_time is now read as UTC, not as the host's local time. The Orq query fields are epoch milliseconds, so the same TraceInput used to select a different window on every machine that ran it, and a naive bound compared against a timezone-aware one raised TypeError inside model validation. Both fields are stored timezone-aware after validation; pass an aware datetime to select any other zone.
  • orq_evaluator's result name now carries the model override. orq:<evaluator_id> when model is omitted, orq:<evaluator_id>@<model> when it is set. Two orq_evaluator calls differing only by model — the standard judge comparison — used to produce two result streams under one indistinguishable name. The evaluator also resolves its Orq client once per factory call instead of once per row, and accepts api_key= and base_url=.
  • A recorded turn that both answered and called a tool now renders both. output_to_text used to return the assistant text alone for an AgentResponse, so a judge asked whether the agent looked something up saw only the prose. Text and tool-call markers now render together through messages_to_text; scores from before this change are not comparable for mixed turns.
  • Classify answers now survive per repetition and appear in pairwise reports. The additive JuryRepetition.raw_output field preserves the complete validated answer in EvaluationResult.raw_output["jury"]; pairwise observations retain each ordering's answer in saved JSON, while expanded HTML/dashboard comparisons show choice, confidence and probabilities in the run's original A/B frame. Missing details remain absent, and older saved records still load. send_results_to_orq() still strips raw_output, so the hosted Orq experiment view does not receive these details.
  • Breaking for custom simulation hooks: on_evaluator_complete(datapoint_id, name, score: float, result) is replaced by on_evaluator_complete(datapoint_id, name, score: EvaluatorScore, result: SimulationResult). Update the parameter names and read the verdict from score.score.value, its explanation from score.score.explanation, and failures from score.error. The hook now fires for errored and non-numeric evaluator outcomes too; numeric values alone enter metadata['evaluator_scores'], unusable outcomes enter metadata['evaluator_errors'], and reports count them as Dropped.
  • Every judged llm_jury() result now carries the full jury record on EvaluationResult.raw_output["jury"], on every assignment — raw_output was None there before. The panel already computed a complete JuryResult (per-judge votes with model ID, verdict, explanation, abstain/failure state, and each judge's raw per-repetition verdicts), and the scorer attached it only under assignment="cyclic"; under the default assignment="all" the whole structure was discarded and the caller kept a one-line summary on explanation naming no judge and no per-judge verdict. Detail that was computed and billed cannot be recovered without re-running the panel, so it is now attached unconditionally rather than behind an opt-in a caller has to know to pass. JuryVote.repetitions is now a list of JuryRepetition objects, each carrying that pass's value and explanation; saved reports with the old bare scalar list still load, with missing explanations as None. A caller reading if result.raw_output is None off a jury evaluator as a signal now takes the other branch — nothing in evaluatorq did, and send_results still strips raw_output before uploading to the Orq platform, so this is local-only data either way. The one case that still carries no record is a datapoint whose target errored: no panel ran, so there is nothing to record, and it returns inconclusive with raw_output=None as before.
  • The jury summary line omits the agreement clause instead of printing raw agreement n/a. A single cyclic vote and a numeric panel both have no cross-judge agreement to report, and a placeholder sitting in a fixed slot reads as a measurement that came back empty rather than one that does not apply. [jury: 3/3 judges, raw agreement 67%] is unchanged; [jury: 1/1 judges, raw agreement n/a] is now [jury: 1/1 judges]. Update any assertion or log scraper matching the literal raw agreement n/a.
  • A panel that ran but reached no verdict now records its judge failure under EvaluationResult.raw_output["evaluation_error"]. The value uses the same RunError shape as red-team evaluation errors and includes the last judge error, all recorded judge errors, and the failed-judge count. The key is absent when judges abstained cleanly or simply failed to reach the quorum without recording an error. A target error still leaves raw_output=None because no panel ran, and the direct llm_jury() path leaves the payload in raw_output for the caller to inspect; only the red-team report converter promotes it to the typed run error rollup.
  • Pairwise JudgeStats.consistency and consistency_raw are now populated under both plurality and bt-sigma. They measure repeated-pass self-consistency directly from the observations and do not require a Bradley-Terry fit; sigma remains available only from the BT-sigma fit. HTML reports show consistency columns when any judge has data, use n/a for judges without measurable repeats, and explain the empty case in a caption.
  • to_open_responses() now preserves simulation state in fixed locations. Its metadata always includes criteria_verified, criteria_meta, criteria_errors, scorer_errors, and datapoint_id, and the response envelope includes token_usage_known beside usage. On the evaluatorq job path, criteria_errors and scorer_errors are None because conversion happens before scoring; an unserialisable metadata value is published as its repr() with a warning.
  • Built-in simulation evaluator detail now has structured local output. criteria_met attaches its criteria records under raw_output['criteria'] with an audited label, and conversation_quality returns a ConversationQualityScore whose breakdown contains component scores and weights. This raw_output is available on the returned EvaluationResult and local result dumps, but is stripped before the Orq experiment upload and is not attached to the evaluator span.
  • Returned simulation results now retain evaluator detail. Each SimulationResult.evaluator_details mapping contains the structured raw_output for evaluators that provide it, including criteria evidence for criteria_met, component scores and weights for conversation_quality, and custom evaluator output. The field is also preserved in saved SimulationRun JSON and JSONL result exports; the mapping is empty when no evaluator provides structured output.
  • on_stage_end() now receives stage failures in meta['error'] as the live exception object, including for generate(). DefaultHooks logs the failure at WARNING level and RichHooks prints a failed-stage line; custom hooks should test whether the key is None before formatting the exception.
  • sim_model= is removed from every simulation entry point; llm_config= carries the model. simulate(), generate_and_simulate(), generate(), generate_personas() / generate_scenarios() (and their singular forms) and extend_from_experiment() no longer accept the keyword — pass llm_config=LLMCallConfig(model=...), which says the same thing and can carry temperature, reasoning_effort, timeout_ms, extra_body and a client beside it. Two spellings of one setting could not report an explicitly-passed default as a contradiction: simulate(sim_model=DEFAULT_MODEL, llm_config=LLMCallConfig(model='other')) ran on other and warned about nothing. Update any call passing sim_model= — there is no deprecation shim. Unaffected: the CLI's --sim-model flag, which now builds that config for you, and the model= argument on SimulationRunner, the generators and the trace helpers, which still folds into a config beside it.
  • JudgeAgent pins itself to the Responses API and logs when a config says otherwise. llm_config.api is one knob for both simulation agents, and only one of them can honour every value: the judge sends function tools and reasoning_effort in one request, which chat completions answers with a 400 on models like gpt-5.4-mini, while the user simulator's plain completion works on either endpoint. Same model, two roles, two viable protocols. LLMCallConfig(api='chat_completions') now applies to the user simulator and is overridden on the judge with a WARNING naming it, where it previously reached the judge and broke the run. The pin is JudgeAgent.REQUIRED_API; the general default is still BaseAgent.DEFAULT_API.
  • Every simulation agent defaults to the Responses API when its LLMCallConfig leaves api unset, overriding that class's own chat_completions default. JudgeAgentConfig and UserSimulatorAgentConfig supplied this before; a caller handing an agent a bare LLMCallConfig used to get chat completions instead, which returns a 400 on models like gpt-5.4-mini the moment the judge sends function tools and reasoning_effort together. The default is BaseAgent.DEFAULT_API — a class attribute, not an environment variable, because it is a per-call setting. Set LLMCallConfig(api='chat_completions') to opt out.
  • A job that returns a top-level error key now marks the row as failed, instead of counting as a clean success. process_job read only name and output from a job's return value, so JobResult.error was set only when the job raised. A job that caught its own failure and reported it — which is what the simulation jobs do, because one dead row must not kill the batch — came back with error=None, and a run whose conversation never happened printed Failed Jobs 0, Success Rate 100% and exited 0. Both simulation jobs — simulate()'s own and wrap_simulation_agent's — now emit error unconditionally (None on success, the runner's reason when the run ended in error or timeout), and process_job honours the key while keeping the output, so the transcript survives for diagnosis — its evaluators are skipped, though, because scoring a conversation already known to be dead buys nothing and costs an LLM judge call per row. wrap_langchain_agent now catches a failing invoke/ainvoke and reports it the same way rather than letting it raise; a caller mistake (no prompt and no usable messages column) still raises. The job's OTel span is now marked ERROR on every failed row, including the two raise paths, which previously closed as OK. Presence is the failure signal, not the message's truthiness: a payload that flattens to nothing is reported as job reported a failure with no readable message and logged, rather than silently becoming a clean row. check_pass_failures(results, treat_errors_as_failure=True) catches these rows as a result — that flag still defaults to False, so a caller who does not pass it gets the same exit code as before; the change is to the reported counts, not to any exit status. Any other job already returning a top-level error key for a non-failure reason will now be counted as failed.
  • simulate(exit_on_failure=True) now raises when a run ended in error or timeout, not only when a row was dropped. The check counted rows missing from the result cache, and a run the runner ended in error is in the cache — so eq sim and simulate() returned normally over a target that was 401 throughout, which is what the caller asked to exit on. SimulationDroppedError's message now names both counts. exit_on_failure=False is unchanged: the same rows are reported as a WARNING. Scorer verdicts are still reporting only — an agent that answered badly does not raise.
  • LLMCallConfig.temperature has no default — unset means the parameter is not sent, and the provider applies its own. It previously defaulted to 1.0, and evaluatorq's own call sites layered literals of their own on top (0.8 for persona and first-message generation, 0.9 for edge-case scenarios, 0.7 for the executive summary and chat-completions agent calls, 0.3 for trace analysis, 0.0 for the judge). Reasoning-class models reject the parameter outright rather than clamping it — gpt-5.6-luna answers 400 Unsupported parameter: 'temperature' is not supported with this model — so a hardcoded temperature on the default model turned every persona x scenario pair into a failure and simulate() into a RuntimeError. No call site sends a temperature now; a caller who wants one sets LLMCallConfig(temperature=...), and a per-call temperature=None means unset rather than an explicit null. BaseAgent._resolved_temperature lost its fallback argument accordingly and now gates on model_fields_set like its sibling resolvers, so an explicit LLMCallConfig(temperature=None) opts an agent out rather than deferring to a call site. The judge is the one behaviour change worth planning around: its scoring calls are no longer pinned to temperature=0.0, so judgments are no longer reproducible run-to-run by default. That affects every simulate() caller, not only those on a reasoning model. Pass JudgeAgentConfig(temperature=0.0) to restore it — on a model that accepts the parameter.
  • A set-but-empty ORQ_OTEL_* tuning variable now logs a WARNING and falls back to the default, instead of falling back silently. _env_int treated an empty string like an unset variable, so an unresolved workflow variable in a CI env: block expands to the empty string and disabled the knob with no signal. Whitespace-only values are treated the same way. Related: ORQ_OTEL_MAX_BATCH_SIZE larger than ORQ_OTEL_MAX_QUEUE_SIZE is still clamped down to the queue size, but the clamp now announces itself with a WARNING rather than happening silently.
  • EVALUATORQ_REASONING_EFFORT has no default — unset means the parameter is not sent, and the model applies its own. It previously fell back to "medium" for the simulator's own calls (user simulator, judge). A global effort is the wrong default in both directions: on a model that does not accept the parameter it costs a rejected request plus a retry per (model, tool shape) — memoised per process, so a short run or CI job never amortises it — and on a model that does, it silently overrides the provider's own tuned value. Set the env var, or LLMCallConfig.reasoning_effort on the agent's config, when you actually want a specific effort. Simulation only; red teaming's target_reasoning_effort was already opt-in.
  • OrqResponsesTarget.retry_attempts now defaults to 1 — a single attempt, no retry — down from falling through to with_retry's default of 5. common.target_call.call_target_with_retry is the single retry owner for target calls on every surface that drives a target (red team static, hybrid, pipeline, orchestrator, and simulation); a target that also retries internally multiplies against that budget instead of adding to it — 5 inner attempts under 3 outer ones is 15 calls to a target that is already refusing. Raise retry_attempts only when constructing the target directly and calling respond() outside call_target_with_retry.
  • Env-var overrides now share one reader, and a misconfigured EVALUATORQ_LLM_TIMEOUT_S / EVALUATORQ_LLM_MAX_TOKENS warns and falls back to the default instead of raising. common.env_config (env_int / env_float / env_bool) is now the single place env overrides are parsed and validated: unset falls back to the default silently, and a set-but-empty/whitespace, unparseable, out-of-range, or non-finite value logs a WARNING and falls back to the default. It never raises. The private readers in tracing/setup.py and simulation/agents/base.py route through it. This changes the two simulation knobs above, which previously raised a ValueError on a non-numeric value and crashed the process at import (DEFAULT_MAX_TOKENS is computed at module scope); they now warn and use the default, matching the non-fatal contract the ORQ_OTEL_* tracing knobs already followed, and both are now bounded with min_value=1 so 0, a negative, or nan/inf also fall back with a warning rather than reaching the provider. ORQ_DISABLE_TRACING also now recognises yes / on and is case-insensitive, in addition to the previous 1 / true. A shared reader with env_int / env_float existed briefly on the RES-1286 branch and was removed before merge in favour of pydantic Field bounds on the recommendations config (which reject a meaningless value instead of warning and falling back); this reintroduces it deliberately for the process-global tuning knobs, where a warn-and-continue contract is wanted over a hard failure. EVALUATORQ_CAPTURE_MESSAGE_CONTENT and EVALUATORQ_PROPAGATE_TRACE_CONTEXT route through env_bool too, so they now recognise yes/on/no/off case-insensitively and warn on an unrecognised value instead of silently reading it as false — previously anything but true/1 turned the toggle off with no signal. EVALUATORQ_CATALOGUE_TIMEOUT_S routes through env_float(min_value=0.1) and no longer raises at import on a non-numeric value. EVALUATORQ_SPAN_MAX_TEXT_CHARS, EVALUATORQ_REASONING_EFFORT, ORQ_DEBUG and COLUMNS are left as bespoke reads (a capture-all sentinel whose default is None rather than a number, a string enum, a debug-only truthy-any toggle, and a terminal probe respectively).
  • LLMCallConfig.completion_params() and .responses_params() are removed. Both are replaced by a single LLMCallConfig.request_params(*, api=None, **params), which renders the shape the api argument names (self.api when omitted). Two builders meant a call site could render a config that says responses into chat-completions shape and no one would notice — exactly the accepted-then-ignored failure this class exists to prevent. A call site that is structurally single-endpoint passes api= explicitly and gets a warning if that contradicts an explicitly-set self.api. Update any call site using either removed method — there is no deprecation shim.
  • Simulation's Responses calls and the executive-summary narrative now go through the canonical executors, and are priced. BaseAgent._call_responses routes through common.llm_call.execute_response, and generate_executive_summary routes through execute_chat_completion — both now get slot limiting, the reasoning drop-and-retry-once, pipeline metadata, trace headers and, previously missing, a price_usage call. Simulation Responses calls on non-Orq endpoints were previously left unpriced entirely. generate_executive_summary now returns an ExecutiveSummary(text, usage) dataclass instead of str | None — update any caller that unpacked or compared the old return value directly.
  • Post-processing spend now lands in the run totals instead of only the log. Red team's ReportSummary gains post_processing_token_usage (recommendation generation, including trace condensing, plus the executive summary), folded into token_usage_total once both steps have run, been skipped, or failed — previously that spend was log-only and invisible to report.summary. Simulation's new SimulationRun.token_usage_total sums every result's usage plus, for generate_and_simulate(), the GENERATE stage's persona/scenario cost and the executive summary's cost; it does not include recommendation generation, which stays log-only on both surfaces.

  • turn_efficiency no longer pays more for a longer conversation. The scorer's cliffs stopped at 6 turns and the decay past them restarted from 1.0, so a 7-turn run scored 0.9 while a 6-turn run scored 0.7 — the curve went up at the seam. The decay now continues from the last cliff's score (7 turns → 0.6, 8 → 0.5, floor 0.3), so scores fall monotonically with turn count. Only runs that opted turn_efficiency or conversation_quality into evaluator_names are affected — neither is in the default ["goal_achieved", "criteria_met"]. Both scorers' policy is now exposed as simulate(scoring=SimulationScoringConfig(...)) / generate_and_simulate(scoring=...): the cliffs, the decay, the floor and the composite weights, all with the shipped values as defaults. The config rejects unordered cliffs and weights that do not sum to 1.0 at construction.

  • Config that was accepted and silently ignored now takes effect. Several fields validated cleanly and then did nothing, so a caller who set them saw no error and no change. All of them are now honoured, which means runs change behaviour for callers who were already setting them: simulation's BaseAgent read only model/api/client/retry_count from its LLMCallConfig and fell back to literals for everything else, so temperature, max_tokens, timeout_ms, extra_kwargs and reasoning_effort on a JudgeAgent or user simulator were inert; LLMConfig.max_tool_continuations was documented as user config but read from the module-level PIPELINE_CONFIG default, so a tool-using ORQ target still gave up after 5 continuation rounds and reported "unresolved pending tool calls" — which the judge then scored as a genuine weak answer; and red team's static leg dropped target_agent_timeout_ms and max_target_retries entirely (the model/deployment path ran on a hardcoded 300 s ceiling with no retry). Per-call-site literals are now defaults that a caller-supplied value beats, in that order.

  • extra_body is a dedicated parameter, not an extra_kwargs key — and passing it in extra_kwargs now raises. It carries the Orq router body (retry policy, thread and memory ids), which is owned by the call site, so a user routing it through extra_kwargs silently replaced the router body rather than adding to it: retry hints vanished with no signal on the judge, jury and recommendation paths. execute_chat_completion, execute_chat_parse, execute_response and generate_structured all gained extra_body=; check_reserved_keys rejects the key inside extra_kwargs on both endpoints. Update any call site using extra_kwargs={'extra_body': ...} — it now raises ValueError instead of quietly winning. The one merge seam is generate_focus_area_recommendations, where a caller-supplied extra_body merges into the router body.
  • LLMCallConfig.extra_body is the user seam into the request body, and it merges rather than replaces. There are two injection points and they target different parts of the request: extra_kwargs sets top-level SDK call arguments, extra_body sets fields in the HTTP body the SDK has no named parameter for — which is where the Orq router's retry policy and thread/memory ids live. Previously only the call site could write that body, so a user had no way to add a body field at all; the one workaround, extra_kwargs={'extra_body': ...}, replaced the whole thing, silently dropping the router's retry hints. The new field is layered over the call-site body per key: your keys win, the ones you did not set survive. Honoured on both endpoints (LLMCallConfig.request_params(), which renders the shape api names), by the judge, the simulation agents and the Responses target.
  • One reserved-key vocabulary instead of three. contracts.py owns _RESERVED_COMPLETION_KEYS / _RESERVED_RESPONSES_KEYS and the shared check_reserved_keys; common.structured_output no longer keeps its own narrower copy, which had drifted to omit extra_body on the chat legs and to reserve max_output_tokens while leaving max_completion_tokens overridable. The guard also now runs inside the common.llm_call executors, so a caller that reaches one without an LLMCallConfig gets the same protection.
  • reasoning_effort reaches the provider on every role, and on both endpoints for the target, judge and jury. It was read at exactly one call site (OrqResponsesTarget), so setting it on the attacker, the judge, the jury or a simulation agent validated and did nothing. LLMCallConfig.request_params() now renders the flat chat-completions spelling or the Responses reasoning={'effort': ...} block depending on the endpoint (self.api, or an explicit api= at a structurally single-endpoint call site). The attacker role is chat-completions only: its call sites pass api='chat_completions' explicitly, so attacker=LLMCallConfig(api='responses') does not switch the endpoint. The two sites that build params through request_params() warn when you set it; the rest pass their fields to the executor directly and cannot warn. EvaluatorConfig now subclasses LLMCallConfig and inherits temperature, max_tokens, timeout_ms, extra_kwargs, extra_body, reasoning_effort and client verbatim — it previously hand-listed them and set extra='forbid', so LLMConfig(evaluator=LLMCallConfig(reasoning_effort=...)), the documented pattern, raised a ValidationError. Three fields are still declared on the subclass deliberately: model (nullable judge shorthand, excluded from dumps), api (defaults to responses, not chat_completions) and retry_count (identical declaration, judge-specific description). An import-time guardrail compares each inherited field's annotation, default and constraints against the parent and fails loudly on any other override — or on an allowlist entry whose divergence has disappeared; a plain field-name comparison could not see a default drifting apart, which is the failure that actually occurred. red_team()'s target_reasoning_effort reached one of five target constructions and its preflight could raise ValueError for targets that never receive it; the preflight is now scoped to targets whose backend actually sends it, and the rest warn. simulate() and generate_and_simulate() gained target_reasoning_effort so the two surfaces measure the same agent.
  • The judge's Chat Completions legs now forward reasoning_effort and extra_body too. run_judge passed only extra_kwargs on both of them, so an effort or a router body set on an EvaluatorConfig — or via llm_jury(reasoning_effort=...) / llm_jury_pairwise() / PairwiseComparator, which build their config internally — worked on the Responses judge and was silently inert the moment the call fell back to chat completions (structured_output=False, a client that does not route through the Orq router, or a model the Responses endpoint will not take). Jury verdicts change for callers already passing reasoning_effort=. Red team's recommendation call sites likewise never read EvaluatorConfig.extra_body and never passed reasoning_effort at all; the router body there is now layered call-site-first, so a caller key wins per key.
  • A failed model-catalogue fetch no longer poisons the process. One HTTP hiccup against GET /v2/models cached {} for the whole run, silently degrading every later call to unpriced and chat-completions-only. The fetch is now retried up to three times before the empty result is cached, and its timeout is settable via EVALUATORQ_CATALOGUE_TIMEOUT_S. New register_model() / get_model_info() / clear_model_overrides() make the catalogue writable: a model Orq does not list — a self-hosted deployment, or one newer than the workspace catalogue — could previously neither be priced, qualified for Responses, nor have its reasoning effort validated, and there was no way to see or fix that.
  • register_model() keys on the bare model id, and ModelInfo normalizes an empty reasoning_efforts to None. 'openai/gpt-x' and 'gpt-x' name one model and now register one entry: the provider/ prefix is stripped on write, because every internal lookup asks for the bare spelling — a qualified registration stored verbatim was never found, so the override was a silent no-op. Registering both spellings replaces rather than duplicates. ModelInfo(reasoning_efforts=frozenset()) becomes reasoning_efforts=None ("the catalogue does not say"), the only other reading the field allows: an empty set made validate_reasoning_effort reject every value, defaults included, with an empty accepted-values list.
  • Red team's static leg reports the same error taxonomy as hybrid. It passed neither the run config nor the backend's map_error to its job factory, so an identical target failing identically produced orq.http.429 in hybrid mode and a generic target_error in static. The self-judge / family-bias guard was inert there too — it compared raw target strings like agent:abc123, which resolve to provider family unknown, so judging gpt-5 with gpt-5 warned in dynamic mode and never in static.
  • Previously-hardcoded budgets are now parameters, with their current values as defaults: with_retry's backoff curve (min_wait_s, max_wait_s, jitter_fraction), generate_structured's 300 s timeout, the executive summary's 400-token cap, apply_recommendations' four merge budgets, simulation's 500-char tool-result truncation (which shapes what the simulated user reacts to and the judge then scores), the adversarial-timeout abandon threshold, the objective-generation batch size, and the black-box probe turn budget. The executive summary also stopped raising a swallowed TypeError on extra_kwargs={'temperature': 1} — the very escape hatch its own docstring documents for reasoning-class models — which surfaced as a silently absent summary.
  • New CLI flags. eq redteam run gained --target-timeout-ms, --max-target-retries, --retry-count, --max-tool-continuations and --target-reasoning-effort; eq sim run gained --target-reasoning-effort, --persona-seed and --scenario-seed. simulate() / generate_and_simulate() gained target_agent_timeout_ms, max_target_retries, per_simulation_timeout_s, max_tool_result_chars and edge_case_percentage — the per-simulation wall clock previously existed only on run_batch, which the public path bypasses, so a stalled conversation had no overall bound and TerminatedBy.timeout was unreachable.
  • SimulationConfig's four new numeric knobs are bounded, and per_simulation_timeout_s=0 now raises. target_agent_timeout_ms and max_tool_result_chars require > 0, max_target_retries >= 0, and per_simulation_timeout_s > 0 — None is the only spelling of unbounded. 0 previously validated and then reached a timeout_s <= 0 sentinel that read it as no bound at all, the opposite of what typing it means.
  • EVALUATORQ_LLM_TIMEOUT_S, EVALUATORQ_LLM_MAX_TOKENS and EVALUATORQ_REASONING_EFFORT are resolved at call time, not import time. Setting os.environ[...] after importing evaluatorq was previously a no-op, and the values could not differ per agent or per run. They are now fallback defaults that a config value beats. Still simulation-only.
  • The Orq Responses target honours the config it is handed and survives a model that rejects reasoning. It built its request from four fields and dropped temperature, extra_kwargs, api and retry_count, which meant no Responses option outside that set — top_p, store, truncation, tool_choice, parallel_tool_calls, … — was reachable at all. It also never applied the drop-and-retry-once contract every other Responses call uses, so a target model that 400s on the reasoning block failed every datapoint in the run rather than retrying once without it. Reasoning effort is now recorded as its own span attribute, so two runs at different efforts are distinguishable in traces even with EVALUATORQ_CAPTURE_MESSAGE_CONTENT=false.

  • Structured generation degrades through four chat rungs instead of two, and parses what it gets back. generate_structured now runs strict parse() → non-strict json_schema → a forced tool call carrying the same schema → bare json_object, and every text rung's output goes through the canonical fence-tolerant parser (common.extract_json) plus schema validation before being returned. Callers that used to receive (None, raw_content) on any fallback and fence-strip it themselves now get (parsed_model, raw_content) — three call sites (persona_generator, scenario_generator, traces) never did that stripping and dropped a fenced ``json payload on the floor.parsed is Nonenow uniformly means "no rung produced anything that validated", with the last non-empty text still returned to log. The forced-tool rung is what makes this a *stricter* backup rather than a looser one:tool_choicenaming the function leaves the model no prose channel, and function calling is a different provider capability thanresponse_format, so models that 400 on a JSON schema often still support it. **It is the only rung that edits the prompt** (one appended user turn telling the model the tool will be called) and it is skipped entirely when the caller passed their owntools/tool_choice. **This can cost up to two extra calls** on a provider that answers nothing usable — previously the ladder gave up after two. Truncation and refusals still raise on every rung rather than continuing. The rung that answered is recorded on the span asorq.structured_output.leg; the old booleanorq.structured_output.fallbackis still written. (Return shape: as of the usage change below, the pair is now theparsed/rawfields of aStructuredResult`.)

  • generate_structured returns what the ladder cost, and every rung is counted. It returns a StructuredResult (parsed, raw, usage) instead of a (parsed, raw) tuple — unpacking the result no longer works; use the fields. The dataclass exists because the call can bill up to five provider requests (the Responses leg plus four chat rungs) and none of that usage was extracted at all, so a call that degraded twice was reported as free. usage is the sum over every rung that reached the provider, not just the one that answered, priced per rung through price_usage. A rung whose usage block cannot be read is counted as one unpriced call with a warning naming the cause — never as zero — so priced_calls < calls marks the figure as a lower bound. The eleven call sites (persona/scenario generators, traces.py, and both recommendations.py modules) all sit on paths with no report field for this spend, so each phase logs its own total (Persona/scenario generation: N tokens over M LLM call(s), $X); PersonaGenerator/ScenarioGenerator gained get_usage()/reset_usage() mirroring BaseAgent. Private helpers traces._summarize_conversation and redteam.reports.recommendations._condense_attack now return their usage alongside their result. (RES-1295)
  • generate_structured disarms the client's SDK retry budget before the ladder runs. Every rung is already wrapped in with_retry, so the two layers multiplied — on evaluatorq's own default path, not just on an injected client: resolve_llm_client builds clients with max_retries=2 and red team's recommendation client with retry_count=3, so five outer attempts over two SDK retries was fifteen requests per rung. It now calls without_client_retries once at the top, as common.judge already did. Callers who were relying on the stacking see strictly fewer requests on a 429/5xx storm — raise retry_count to compensate. The clone never mutates the client you passed.
  • A structured-generation call that raises still reports what it billed. New StructuredGenerationError in common.structured_output (subclassing RuntimeError, so an except RuntimeError keeps working) carries the ladder total on .usage; a provider error propagates as itself — masking a 429 behind a RuntimeError would cost the caller the status code — with the same attribute attached in place, so callers harvest it with getattr(exc, 'usage', None) rather than by exception type. Truncation is the common case: LengthFinishReasonError hands back the completion it refused to parse, so a parse() rung that generated a full max_tokens of output is no longer counted as free. The three legs that swallow the exception — trace summarizing, and both recommendations.py modules — now fold that spend into their phase totals, which previously reported "no usage reported by the provider" for an entire failed ladder.
  • Judge verdicts are generated explanation-first, and a self-contradictory abstention is coerced. EvaluatorResponsePayload and the dynamic jury verdict model now declare value last, so the structured-output schema makes the model write its reasoning before committing to a verdict rather than justifying one it already emitted; DEFAULT_SECURITY_EVALUATOR_SYSTEM_PROMPT was reordered to match, since a prompt that lists value first fights its own schema. A payload that arrives with abstain=True and a non-null value now has the value dropped, with a logger.warning, a judge.verdict_coerced span attribute and JudgeOutcome.verdict_coerced. Previously that contradiction flowed on unchecked: the static OWASP and adaptive evaluator paths read payload.value without consulting abstain at all, so a judge that declined to answer was still scored as a decisive pass or fail, and any vote built straight from the pair would trip JuryVote's validator after the judge had been billed. The system prompt now also describes abstain — it previously named only two keys while the enforced schema had three. The mirror shape (value=None, abstain=False) is unchanged: the redteam paths still count it as a failed repetition, llm_jury still reads it as an abstention.
  • The simulation pipeline runs its own LLM calls on the Responses API. UserSimulatorAgentConfig now defaults api='responses' (the judge already did), and the persona, scenario and first-message generators call the Responses endpoint, so a run's spans are uniformly responses {model} instead of a mix the trace UI types differently — which is what made the simulator and generator calls look absent when scanning for router spans. Pass api='chat_completions' on the config to opt out. generate_structured gains api= with the default unchanged, so the red-team report paths are unaffected; on api='responses' it degrades to the chat legs, with a warning naming the cause, when the endpoint is absent (404), when a 400's body names the schema form as unsupported, or when nothing comes back parsed. Any other 400 raises rather than being blamed on the provider. The user simulator no longer sees the target's tool calls or tool results — after role inversion those rows are a user message carrying tool_calls followed by orphan tool rows, which the provider rejects, so every simulation against a tool-using agent died at turn 1. The canonical transcript is untouched: the target and the judge still see the tool traffic.
  • Scenario and persona generation size their max_tokens from the count they were asked for, instead of a flat cap. A batched structured call for ~15+ items truncated, and truncated structured output is unrecoverable (both legs raise rather than retry at the same budget). This raises spend for large counts — a 30-scenario request now asks for 15 000 output tokens where it asked for 6 000. The budget is deliberately unbounded at the top: a count large enough to exceed the model's own limit fails with the provider naming the limit, which beats silent truncation at an arbitrary ceiling.
  • One default model for every surface: openai/gpt-5.6-luna. evaluatorq.contracts.DEFAULT_PIPELINE_MODEL is now the single source; simulation.types.DEFAULT_MODEL, llm_jury.DEFAULT_JUDGE_MODEL and the dashboard's DEFAULT_APPLY_MODEL are aliases of it rather than three independently drifting literals (previously gpt-5-mini, openai/gpt-5.4-mini, openai/gpt-5.4-mini and gpt-5.6-luna). Each surface stays individually overridable — --attack-model / --evaluator-model, --sim-model, judges= / model=, EVALUATORQ_APPLY_MODEL — only the fallback is shared. Every default is now provider-prefixed for the Orq router; targeting OpenAI directly (only OPENAI_API_KEY set) means passing the bare gpt-5.6-luna, which the red-team defaults did not previously require.
  • DEFAULT_TARGET_MAX_TOKENS is 10_000, up from 5000, and simulation now shares it. On a reasoning model the budget covers hidden reasoning as well as the visible answer, so the old figure left much less room for output than it used to — truncation surfaces as a judge PARSE error (incomplete_details.reason=max_output_tokens) or, in simulation, finish_reason=length before any tool call. It caps LLMConfig.max_tokens (attacker), EvaluatorConfig.max_tokens (judges), OpenAIModelTarget, create_deployment_job, and now simulation's DEFAULT_MAX_TOKENS, whose EVALUATORQ_LLM_MAX_TOKENS override is unchanged (its default was a separate 8192).
  • The dashboard's apply-recommendations merge no longer has its own default model. It was deliberately set apart from the pipeline's — rewriting production agent instructions warrants a stronger model than scoring does — and now shares the fallback like every other surface. EVALUATORQ_APPLY_MODEL is how you raise it again, unchanged.
  • Multi-turn requests now carry Anthropic prompt-cache breakpoints, on Orq-routed clients only. Simulation agents and the red-team OpenAIModelTarget mark the system prompt and the end of the persisted transcript with cache_control: {"type": "ephemeral"} — on the Chat Completions path as a marked content block, on the Responses path as a marked input content part. Both are gated on caching_applies(client, model), which requires the Orq router and an Anthropic (or agent/) model: cache_control inside a content part is outside the direct OpenAI schema, so a client pointed at api.openai.com or a self-hosted OpenAI-compatible endpoint is left untouched, and a routed non-Anthropic model gets no needless request-shape change (Orq documents the marker as ignored, not rejected, there). Anthropic caching is opt-in — without a breakpoint an append-only transcript is re-encoded in full every turn — while OpenAI, Gemini, DeepSeek and xAI cache automatically, so no prompt_cache_key is set and no per-provider branch exists. On a routed client, marked content is sent as a text-block list rather than a plain string; a target that inspects the raw request body will see that shape. A cache write costs 1.25x, so neither helper marks an input whose text is below CACHE_MIN_PROMPT_TOKENS (1024) — no Anthropic model caches below that, and the write would be pure loss. Covered: the simulation judge and user simulator, the red-team OpenAIModelTarget, and OrqResponsesTarget — which is both the simulation agent:<key> target and the execution half of the default red-team agent backend. Not covered: redteam/runtime/jobs.py and common/judge.py, which replay conversations through execute_chat_completion without breakpoints (RES-1360). Callers say how many trailing messages they rebuild each turn via the required volatile_tail keyword — marking a message that does not persist pays a 1.25x write nothing ever reads back, which is what the simulation judge's per-turn instruction did. Both paths are confirmed against live traces on anthropic/claude-sonnet-4-6 with a uuid-salted cold prefix, three judgements each: Chat Completions read 0 / 6,991 / 7,881 tokens, Responses 0 / 7,417 / 8,304 (higher because its system prompt lives in instructions, outside the marked input).
  • The adaptive red-team attacker loop is now cached too. MultiTurnOrchestrator.run_attack marks the adversarial system prompt and the end of the persisted transcript on every turn, gated on caching_applies exactly like the other surfaces. The call sits inside the turn loop deliberately: apply_cache_breakpoints returns a copy and the transcript grows each turn, so a hoisted call would freeze the turn-1 snapshot and send it for the rest of the attack — a silent correctness bug, not a cost one, which is why the call-site test pins placement rather than just the marker count. Measured live on anthropic/claude-sonnet-4-6 over five turns with a uuid-salted cold prefix: reads of 0 / 1,986 / 2,211 / 2,435 / 2,696 tokens against inputs of 1,988 / 2,214 / 2,438 / 2,699 / 2,948 — a 91.5% hit rate by turn 5, with only the newly appended pair paying full price. No public surface changes; an operator running the default openai/gpt-5.6-luna attacker sees no difference, since OpenAI caches on prefix without a marker.
  • A caller that rebuilds its trailing message every turn must say so. apply_cache_breakpoints(messages, volatile_tail=N) keeps the breakpoint off the last N messages: mark one and the next turn puts transcript content at that position, the prefix diverges right after the system message, and the whole transcript pays a write that is never read. The simulation judge appends a per-turn instruction and passes volatile_tail=1; so does the user simulator's first-message call. The Responses path takes the same guarantee through mark_responses_input(input, volatile_items=N) — items, not messages, because one Message with tool calls renders to several input items; responses_volatile_items(messages, volatile_tail=N) converts between the two and is the only supported way to do so.
  • The simulation judge's system prompt is now byte-stable across a conversation. The [ALREADY CONFIRMED] markers for settled criteria moved out of the criteria listing and into the trailing user message that JudgeAgent.evaluate appends. The markers sat at token position 0, so every mark_settled call invalidated the cached prefix — system prompt, whole transcript and tool schemas — to save a few dozen output tokens. The instruction is unchanged in substance and a custom judge is unaffected, but it now reaches the model later in the request; coverage is structural — no live-model test exercises the new placement.
  • red_team() takes recommendations= instead of generate_recommendations=, and simulation gained the same flag with the same three forms: True (defaults), False (skip the LLM call), or a config instance — RedTeamRecommendationConfig / SimulationRecommendationConfig, both bounded and extra='forbid'. Two removals with no deprecation shim: generate_recommendations= on red_team(), and the long-deprecated config= alias for llm_config= (the shim had targeted removal in 1.4.0 and outlived it). generate_focus_area_recommendations() likewise takes recommendations= in place of its max_areas / max_traces pair. Rename the keyword at the call site; behaviour is unchanged. (RES-1286)
  • Simulation now generates remediation suggestions in-run rather than from a CLI post-run hook, so a saved run carries them regardless of caller. simulate() / generate_and_simulate() default the flag to False, because the returned SimulationResult list has nowhere to carry suggestions — with save/report both unset the run would pay for them and drop them, which now logs a warning. eq sim run and eq redteam run keep --recommendations on by default. (RES-1286)
  • Assistant turns in a replayed transcript are sent to the Responses API as output_text parts. A bare string — or a list of input_text parts — under role: "assistant" is silently dropped by the Orq router (some backends 400 instead), so every stateless Responses target and the simulation judge/user-simulator were replaying history with the agent's own turns missing. The simulation judge saw a transcript with no agent replies and reported "the agent has not yet responded", which is the deeper reason no criterion about agent behaviour could ever fail (RES-1308). Affects OrqResponsesTarget, OpenAIAgentTarget, red-team multi-turn replay, and simulation. An image part on an assistant turn is not representable and is now dropped with a warning.
  • The criteria_met scorer returns 0.0 (was 1.0) for a simulation that ended in an error or a timeout, and logs a warning. Such a run terminates before the judge audits anything, so its criteria outcome is unknown — scoring it a perfect 1.0 let a dead target inflate the run average and conversation_quality. A run with no criteria at all still scores 1.0.
  • Simulation Scenario criteria are now scored from an explicit per-criterion audit the judge returns on every turn (Judgment.criteria_verdicts), folded across the whole conversation. Previously pass/fail was inferred from the absence of a criterion id in rules_broken, so must_happen criteria could never fail and criteria_met returned 1.0 on every run (RES-1308). This changes scores for existing callers who use criteria: criteria_met, conversation_quality, rules_broken and criteria_results can now report failures where they previously reported none. A must_happen criterion passes if it occurred in any turn; a must_not_happen criterion fails if it was violated in any turn. A custom judge that does not emit criteria_verdicts falls back to the old behaviour, logs a warning naming the scenario, and is marked SimulationResult.criteria_verified = False. Judgment.criteria_verdicts is list[CriterionVerdict] | None. CriterionVerdict (new, public, in evaluatorq.simulation.types, re-exported from evaluatorq.simulation) reports criterion_id, occurred and evidence for one criterion on one turn — occurrence only, never pass/fail. None means the judge reported nothing (unknown, criteria_verified=False); [] means it audited and had nothing left to report; a non-empty list is evidence. New public helpers criterion_id_for(index) and CRITERION_ID_PATTERN (both also re-exported from evaluatorq.simulation) fix the criteria_N id format in one place.
  • The simulation judge's finish_conversation tool no longer takes a rules_broken argument. Violations are derived in code from the occurrence audit and Criterion.type; the free-text list is the channel that could not fail a must_happen criterion in the first place, and asking for both gave them something to disagree about. A criterion the audit skipped now keeps its not-observed default instead of being rescued from free text. Judgment.rules_broken and SimulationResult.rules_broken are unchanged as outputs — only the tool input is gone, so a custom judge that populates the field itself still works.
  • The judge stops re-auditing a criterion once it is confirmed to have occurred. Occurrence is sticky, so a settled criterion cannot change; it stays in the prompt (the judge needs it to decide whether to end the conversation early) but drops out of the per-turn criteria_verdicts payload, which costs an id, a boolean and an evidence quote per criterion per turn. A custom judge without a mark_settled method keeps auditing everything.
  • metadata['criteria_meta'] entries gain audited — whether the judge actually returned an occurrence verdict for that criterion, as opposed to it falling to the not-observed default. A must_happen the judge confirmed never occurred and one it silently skipped both report passed: False; only this field separates them. None for runs saved before the field existed.
  • metadata['criteria_meta'] entries also gain evidence — the quote from the turn where the criterion's occurrence first flipped, sourced from the judge's criteria_verdicts audit. '' when the criterion never occurred, None when no tracker was available (same convention as audited).
  • New SimulationResult.criteria_verified field, and criteria_met returns 0.0 for a run where it is False. It is False whenever the judge returned no per-criterion occurrence audit for any turn — a custom judge predating criteria_verdicts, or the built-in JudgeAgent terminating for safety after an unparseable tool call. Those verdicts came from the free-text rules_broken list, which cannot fail a must_happen criterion, so an all-green result there is unknown rather than passing; scoring it 1.0 reproduced RES-1308 one layer up, with a log line as the only signal. None on runs saved before the field existed, and those keep their previous score.
  • The criteria_met evaluator now reports pass=False — with an explanation naming the cause — on exactly the runs it scores 0.0: one that ended in an error or a timeout, and one with criteria_verified = False. The flag was previously derived from criteria_meta alone, so an unaudited run landed on the evaluator trace span and the uploaded Orq experiment as a green PASS beside its own 0.0. An errored run, which has no criteria_meta at all, reported "No criteria defined for this scenario." and pass=True for a scenario that does have criteria.
  • criteria_met no longer counts an individual criterion the judge never audited as met, on any surface. The scorer reads metadata['criteria_meta'] when present (the only place audited survives — criteria_results is keyed by description and carries no provenance) and counts a criterion only when it passed and was audited, logging a warning naming how many were not; the evaluator explanation prints UNKNOWN [required]: … (not audited) instead of PASS for it and excludes it from pass. This lowers criteria_met for runs where the judge audited some criteria and skipped others — previously the score counted the skipped ones as met while the report's own "N/M criteria met" tally did not, so the two contradicted each other. A criterion the judge settled early is audited (a verdict is what settles it), so mark_settled never costs a run a point; audited: None (a run saved before the field existed) still counts as met.
  • audited and evidence now reach the reports. CriteriaRow (in evaluatorq.simulation.types, the per-criterion view model behind the report sections) gains audited, evidence and a computed state of pass / fail / unknown, and SimulationEntry gains criteria_verified. Every surface renders state, not passed: a criterion that passed only because the judge never audited it shows as not audited (a neutral ?) in the dashboard, the HTML report and the markdown export, is excluded from the "N/M criteria met" tally, and a run with criteria_verified = False says so above the criteria list instead of showing a tally that contradicts its criteria_met score of 0.0. The judge's evidence quote is shown beside the criterion it justifies.
  • A simulation whose target fails mid-run now keeps the criteria audit collected before the failure. The error result carries the folded rules_broken, criteria_results, criteria_meta and criteria_verified instead of rules_broken=[] and no metadata, so a must_not_happen violation the judge confirmed on turn 2 survives the target dying on turn 4 — it previously vanished from the result and the report. (It never reached find_triggers: that helper returns [] for any errored result before it looks at criteria, and still does.) On this path only confirmed occurrence is knowledge: a must_not_happen the judge saw violated stays failed, while a must_happen that had not occurred yet is reported as unknown (row state unknown, audited: False), never as failed — the run was cut short before that criterion had its chance, so folding the not-observed default would invent a failure the judge never made and add a phantom row to the cross-run failure-mode table. A target that dies before the judge audits anything reports every criterion that way, plus criteria_verified=False. Such a run is still scored 0.0 by criteria_met, because it terminated by error.
  • EVALUATORQ_SPAN_MAX_TEXT_CHARS defaults to capturing all message content (no truncation), in both the Python and TypeScript tracing layers. Set the env var to a positive integer (canonical: 8192) to cap span text at that many characters (marker ... [truncated]); -1, 0, or unset all mean capture all. The cap applies uniformly to input and output message content. (RES-715 introduced an 8192 default; RES-899 reverts to capture-all and unifies the TS path, which previously hardcoded a separate 2000-char cap.)
  • evaluatorq() defaults to datapoint_parallelism=10 (previously 1). Evaluations are almost entirely provider-bound I/O, so the old default made the common case pay a latency penalty to protect the uncommon one. This changes behavior for existing callers who omit it: ten datapoints now run concurrently. Pass datapoint_parallelism=1 to restore serial execution — do so if your provider rate-limits at low concurrency, or if your jobs mutate shared state that was previously serialized by accident rather than by design. Red teaming already defaulted to 10; simulation is raised from 5 to 10 in the same release, so the number now means the same thing on every entry point.
  • A target call makes exactly max_target_retries + 1 HTTP attempts, on every path. call_target_with_retry owns target retries, so the SDK's own budget is disarmed at that boundary (without_client_retries, a with_options clone — an injected client is never mutated and keeps its transport, auth, base URL, headers and timeout). Previously the two layers stacked and multiplied: an injected OpenAI or Responses client left at the SDK default made 9 HTTP calls where 3 were intended, at 3× the cost and latency, and a caller who set retry_count on such a client had no way to see it. Judge and pipeline calls are unchanged — each already had a single owner. Simulation agent calls (the user simulator and the judge agent) now honour LLMCallConfig.retry_count: SimulationAgent._call_chat_completions / _call_responses passed no budget to with_retry, so they always used the module default of 5 transport attempts regardless of configuration. They now make exactly retry_count + 1 (default 2, previously 5). The chat-completions path additionally retries once within an attempt on an empty response — a content-level retry, so a model that keeps returning nothing still costs up to 2 × (retry_count + 1) calls. retry_count / retry_on_codes passed for a target call are now ignored with a warning naming the owner rather than silently.
  • evaluatorq()'s datapoint_parallelism now bounds evaluator fan-out too. Evaluators within a job previously ran with unbounded concurrency (a datapoint with 50 evaluators issued 50 concurrent provider calls no matter what the datapoint count said); they now share the same per-datapoint semaphore the jobs use. This lowers throughput for callers who relied on the unbounded behaviour — raise datapoint_parallelism to restore it. The budget is shared, not split: a job releases its slot before its evaluators take theirs, so the two never contend and datapoint_parallelism=1 cannot deadlock.
  • New llm_parallelism= on evaluatorq(), red_team(), simulate(), generate_and_simulate() and generate() — a ceiling on in-flight LLM requests for the whole run, counted per request rather than per task. Defaults to 10 when unset, and that default also covers LLM calls made outside any entry point (e.g. a standalone run_pairwise() or summarize_conversations()), so an unconfigured run cannot flood a provider. This lowers throughput for callers who relied on unbounded fan-out — pass a larger llm_parallelism (or -1 for no ceiling at all) or wrap the call in evaluatorq.llm_concurrency_limit(n) to raise it. This is the knob to size against a provider concurrency limit: datapoint_parallelism bounds tasks, and the task bounds nest (datapoints × jobs/evaluators × jury width), so datapoint_parallelism=10 can mean anywhere from 10 to several hundred concurrent requests depending on the fan-out — the number was never something you could compute a request rate from. Requests routed through common.llm_call (judges, juries, simulation agents, the red-team pipeline, the OpenAI backend) take a slot automatically; a job that calls a provider SDK directly is invisible unless you wrap it in the new evaluatorq.llm_slot() context manager, which is also what closes the gap for the ORQ and LangChain targets. Note this is a concurrency bound, not a rate limit: N slots is N / latency requests per second, so a provider that gets faster raises your request rate at a fixed N.
  • Nested LLM limits now stack. An inner llm_concurrency_limit(20) inside a run capped at 10 stays under 10; an inner limit of 5 lowers the cap to 5. llm_concurrency_limit(-1) adds no cap and cannot lift an enclosing one. At the top level, -1 still disables the default ceiling of 10. Unconfigured runs in the same event loop share the default budget, so their combined load stays under 10; explicitly configured runs have separate budgets.
  • parallelism= is renamed datapoint_parallelism= on evaluatorq(), red_team(), simulate(), generate_and_simulate() and generate(), and --parallelism is renamed --datapoint-parallelism on eq redteam and eq simulate. Both old names still work and emit a DeprecationWarning; EvaluatorParams accepts either field name. With two concurrency knobs the bare name no longer said which one it meant — one counts datapoints, the other counts LLM requests. The OTel span attributes orq.redteam.parallelism and orq.simulation.parallelism keep their keys, so existing trace queries and saved dashboard filters are unaffected. Breaking for hook implementors: the red-team ConfirmPayload and the simulation SimulationRunMeta key is now datapoint_parallelism — a hook reading payload['parallelism'] will KeyError.
  • New --llm-parallelism flag on eq redteam, eq simulate and eq simulate from-traces, exposing the llm_parallelism= ceiling to CLI callers. -1 disables the ceiling.
  • First-message generation fails the datapoint instead of inventing an opening. FirstMessageGenerator.generate() used to return a canned "Hi, I need help with: <goal>" when the model returned nothing usable or a 429/5xx outlasted retries, and that line was then simulated and judged as if it were real. It now retries an empty or refused reply up to 3 times, then raises FirstMessageGenerationError (truncation raises at once, since the same budget would truncate again), and lets a provider error that outlasted retries propagate. DatapointGenerator and simulate() drop only the failed persona x scenario pair for unusable output or an exhausted transient provider error, log a Generated N of M datapoints warning, and raise when every pair fails. Authentication, bad-request, and unexpected errors abort the batch so a configuration mistake cannot produce an incomplete dataset. datapoints_from_traces() still falls back to the trace's real recorded opening. Both simulate() and direct DatapointGenerator calls batch first-message tasks at twice the active LLM ceiling, bounding pending tasks without adding another provider-request cap; a top-level -1 leaves that fan-out unbounded.
  • Red-team reports show run warnings. RedTeamReport.pipeline_warnings used to reach only the console. The HTML report, the Markdown report and the dashboard now show them in a "Run Warnings" section straight after the summary. A category whose strategy generation failed but which still ran its hardcoded strategies now gets a warning too, where before only a category left with zero strategies did.
  • Breaking: per-call-site concurrency knobs are removed, because the llm_parallelism ceiling now bounds those calls. Passing any of these raises TypeError:
  • max_concurrency= on run_pairwise(), llm_jury_pairwise(), PairwiseComparator and run_jury().
  • rate_limit_delay= and max_concurrent_calls= on DatapointGenerator.
  • generation_parallelism= on plan_strategies_for_categories() / plan_strategies_for_vulnerabilities(), and datapoint_parallelism= on generate_dynamic_datapoints() / generate_dynamic_datapoints_for_vulnerabilities().

Delete the argument. To bound these calls, set llm_parallelism= on the entry point you run, or wrap a standalone call in async with evaluatorq.llm_concurrency_limit(n):. The paths these knobs capped (trace summarization and inference at 5, first-message generation at 5 plus a 100 ms delay, strategy generation at datapoint_parallelism) now share the default ceiling of 10. - Simulation's datapoint_parallelism defaults to 10 (previously 5), matching evaluatorq() and red teaming. This raises concurrency for callers who omit it; pass datapoint_parallelism=5 to keep the old value, or set llm_parallelism= to bound provider load directly. Applies to simulate(), generate_and_simulate(), SimulationConfig and eq simulate / eq simulate run. - evaluatorq() never exits the process when an evaluator reports pass_=False; it returns the results so library callers can inspect pass_ and choose their own gate. Red-team and simulation surfaces retain their own explicit failure gates. - loguru is now a core dependency (previously gated behind the [redteam] extra). This slightly widens the install footprint for non-redteam consumers but unifies the logging stack across the package. - openai (>=1.92.0) is now a core dependency (previously gated behind the [redteam] extra). The new llm_jury() evaluator imports it at package load, so every base install pulls it; this widens the base footprint for users who only call evaluate(), in exchange for llm_jury() working without an extra. - datapoints_from_traces() and extend_from_traces() now summarize every trace conversation unconditionally before the persona/scenario or traffic-profile call reads it — a short trace that previously skipped straight to that call now costs one extra LLM call in direct mode too. TraceAnalysisConfig.summarize_above_chars is removed with no deprecation shim; because TraceAnalysisConfig is extra='forbid', TraceAnalysisConfig(summarize_above_chars=...) now raises ValidationError — drop the field. A new summarize_conversations() entry point runs that summarize step directly: call it once and pass the result as summaries= to either function so a run that calls both does not summarize the same trace twice; a summaries= mapping is authoritative, so a trace_id absent from it (because it failed to summarize) is dropped rather than retried. (RES-1286)

  • W3C trace context propagation is now toggleable, and the Orq agent and deployment targets propagate it. EVALUATORQ_PROPAGATE_TRACE_CONTEXT=false (or 0) stops evaluatorq sending traceparent/tracestate on its outgoing LLM, Responses and target calls; the default stays on. Two target legs previously sent no trace context at all, so their server-side execution started its own root trace instead of nesting under the target-call span: the red-team ORQAgentTarget (agents.responses.create) and both deployment legs (deployment() / invoke(), used by the simulation adapter, and the red-team static deployment job). All three now pass the headers via the Orq SDK's http_headers. (RES-1422)

  • llm_jury() and llm_jury_pairwise() now accept prompt and criteria together; the rule was exactly one of the two. It is now at least one: a panel with neither still raises. This exists because the two arguments no longer feed the same judge — an LLM judge reads the rendered prompt (which can pull the rubric in with a {{criteria}} placeholder), while a classify judge is handed criteria directly as its question and never sees a prompt, so a mixed panel needs both. Nothing that was valid before is rejected now. Setting both on a panel that seats no classify judge and whose prompt never renders {{criteria}} logs one WARNING naming that nothing reads the rubric, rather than raising.

Breaking Changes

  • The deprecated Streamlit dashboards are removed: eq redteam ui and eq sim ui no longer exist. Both have printed Warning: eq … ui is deprecated; use eq dashboard instead. since July 2026; eq dashboard is the replacement and reads the same saved run files, with no migration beyond the command name. The redteam and simulation extras no longer pull in streamlit, plotly or watchdog — they now install only huggingface-hub plus chart rendering (vl-convert-python) and, for simulation, orq-ai-sdk. The evaluatorq.redteam.ui, evaluatorq.simulation.ui and evaluatorq.common.ui packages are gone, including simulation.ui.token_display; nothing outside the deleted viewers imported them.
  • red_team() parameter renamed: config= → llm_config=. The old config= keyword was kept as a deprecated alias that emitted a DeprecationWarning; it has since been removed (see the recommendations= entry above).
  • LLMConfig flat fields removed: attack_model, evaluator_model, adversarial_temperature, adversarial_max_tokens, llm_call_timeout_ms, llm_kwargs — replaced by role-based attacker / evaluator sub-configs (LLMCallConfig)
  • wrap_simulation_agent() no longer accepts the evaluators= kwarg. Evaluators are wired through evaluatorq() directly (the framework that consumes the job); callers passing evaluators=[...] will now get a TypeError and should move the list onto their evaluatorq(..., evaluators=...) call instead (RES-594).
  • simulate() and generate_and_simulate() no longer accept agent_key=. The single target= parameter now selects the target: "agent:<key>" or a bare "<key>" (hosted Orq agent via the Responses router), "deployment:<key>" (legacy deployment), an AgentTarget, or a callable. Callers passing agent_key=... get a TypeError; migrate to target="deployment:<key>" (or target="agent:<key>"). The eq sim simulate / eq sim run CLI drops its matching --agent-key flag — use --target deployment:<key>.
  • simulate() and generate_and_simulate() now default upload_results=True. With the move to evaluatorq-native execution the framework's upload is the canonical persistence path — the previous False default left runs with no record anywhere. Set upload_results=False explicitly to suppress (RES-594).
  • AttackerResponse is removed — attacker LLM output is unified onto evaluatorq.contracts.AgentResponse, the same shape targets and simulation agents already return. generated_prompt is now AgentResponse.text, and truncated is derivable from finish_reason == 'length'. Removed outright rather than left as a silent alias: AttackerResponse = AgentResponse would have made AttackerResponse(generated_prompt=...) quietly drop the prompt (AgentResponse ignores unknown kwargs) and would have collapsed isinstance discrimination between the two. Construct AgentResponse(text=...) directly (RES-883).

Migration:

# Before
red_team(target, config=LLMConfig(attack_model="gpt-4o", evaluator_model="gpt-4o-mini"))

# After
from evaluatorq.redteam.contracts import LLMCallConfig, LLMConfig

red_team(
    target,
    llm_config=LLMConfig(
        attacker=LLMCallConfig(model="gpt-4o"),
        evaluator=LLMCallConfig(model="gpt-4o-mini"),
    ),
)
  • AgentTarget relocated: moved from evaluatorq.redteam.backends.base to evaluatorq.contracts. Importing it from the old path now raises ImportError. The Backend ABC stays in evaluatorq.redteam.backends.base. AgentContext, ToolInfo, MemoryStoreInfo, and KnowledgeBaseInfo also moved to evaluatorq.contracts, but — unlike AgentTarget — their old import path evaluatorq.redteam.contracts still works (re-exported, same class objects, isinstance unaffected). Only AgentTarget's old path is a hard break.

Migration:

# Before
from evaluatorq.redteam.backends.base import AgentTarget

# After
from evaluatorq.contracts import AgentTarget
  • AgentTarget unified on respond(messages): respond(messages: list[Message]) -> AgentResponse is now the abstract method every target implements. send_prompt(prompt: str) -> AgentResponse is retained as a concrete back-compat shim on the ABC — it wraps the prompt in a single user message and calls respond. Custom targets that previously implemented only send_prompt must implement respond instead.

Migration (bare custom subclass):

# Before — only send_prompt was abstract
from evaluatorq.contracts import AgentResponse, AgentTarget


class MyTarget(AgentTarget):
    async def send_prompt(self, prompt: str) -> AgentResponse:
        return AgentResponse(text=await my_llm_call(prompt))

    def new(self) -> "MyTarget":
        return MyTarget()


# After — respond is the abstract method; send_prompt is a free shim on the ABC
from evaluatorq.contracts import AgentResponse, AgentTarget, Message


class MyTarget(AgentTarget):
    async def respond(self, messages: list[Message]) -> AgentResponse:
        prompt = messages[-1].content or ""
        return AgentResponse(text=await my_llm_call(prompt))

    def new(self) -> "MyTarget":
        return MyTarget()
- OrqResponsesTarget is now stateless: __call__, _previous_response_id threading, _accumulated_usage, and get_usage() are removed. Conversation continuity is the caller's responsibility — pass the full transcript to respond each turn. Pass the target to simulate(target=...) (auto-routes to the target-agent path) or simulate(target_agent=...) instead of relying on __call__. Per-call token usage is reported on the returned AgentResponse.usage. - ORQAgentTarget last-user contract: respond(messages) forwards only the last user message to the ORQ agents endpoint (server-side state is held via task_id) and raises ValueError if messages[-1].role != "user". The endpoint, task_id threading, and usage accumulation are unchanged. - ChatMessage alias removed: the RES-596 deprecated alias ChatMessage = Message is gone. Import Message from evaluatorq.contracts (the public evaluatorq.simulation.ChatMessage re-export is also removed). - Simulation TargetAgent Protocol removed: the simulation runner consumes the canonical AgentTarget ABC from evaluatorq.contracts. The evaluatorq.simulation.TargetAgent / evaluatorq.simulation.runner.TargetAgent exports are replaced by AgentTarget.

Migration:

# Before
from evaluatorq.simulation.types import ChatMessage
from evaluatorq.simulation import TargetAgent

# After
from evaluatorq.contracts import Message      # ChatMessage was an alias of Message
from evaluatorq.contracts import AgentTarget   # replaces the simulation TargetAgent Protocol
  • CallableTarget forwards the full transcript: the wrapped callable now receives the entire conversation as a list[Message] (previously only the last user turn as a str), so stateless callables retain context across multi-turn attacks. The callable signature changes from (prompt: str) to (messages: list[Message]), and usage_fn from (prompt: str, response: str) to (messages: list[Message], response: str). The former last-turn-must-be-user guard is dropped (matching the other stateless targets). Callables that need OpenAI chat-completion dicts can call Message.to_chat_completion() per element.

Migration:

from evaluatorq.contracts import Message
from evaluatorq.integrations.callable_integration import CallableTarget

# Before
target = CallableTarget(lambda prompt: my_agent(prompt))

# After — read the last turn off the transcript
target = CallableTarget(lambda messages: my_agent(messages[-1].content or ""))

New Features

  • eq redteam run, eq sim simulate and eq sim run take structured input and give structured output. --config PATH (or - for stdin) carries the keyword arguments of red_team(), simulate() or generate_and_simulate() as JSON, which reaches every data-shaped parameter, including ones with no flag such as inline personas and scenarios, attack_techniques and edge_case_percentage. Unknown keys are rejected at every depth. A flag passed on the command line beats the file, field by field, and a field neither sets takes the SDK's default. --llm-config '<json>' merges into the LLM configuration, and the narrow model flags win over it for their own field. --json prints the final RedTeamReport or SimulationRun on stdout and moves all other output to stderr, keeping exit codes. eq redteam schema and eq sim schema print the JSON schema of the config file (--input, the default, with the CLI's own defaults) or of the result (--output). YAML is not accepted.
  • New evaluatorq.signals package — computes deterministic ADR-25 structure, tool, autonomy and trajectory-tag signals from ATIF trajectories, with evidence, preconditions and approximation status; includes an evaluator adapter and opt-in Jev tool-role classification. Bundled tag thresholds are calibrated on local Claude Code sessions.
  • Added the trace finder surface. eq find and the dashboard's Trace search page compile a natural-language question, select metadata filters with the classify model, search recent Orq traces through OQL, project each conversation, and classify matches with the classify model. Review-first mode, live progress, trace inspection, and JSON export are included.
  • Classify judges on a jury panel — llm_jury() and llm_jury_pairwise() can now seat the Orq classify model typesafe/jev-latest (Jev) beside prompted LLM judges. A classify judge is not prompted: it is handed the material to judge and one question about it, and answers with a probability distribution — about $0.042 per million input tokens with output free, in under a second. criteria is its question and is required on a panel that seats one; labels={label: description} builds a choice question, levels=[...] (2 to 10 ordered descriptions, numeric mode) builds a score question, and boolean mode compares the returned probability against threshold. state_fields=[...] picks which template paths are handed over as the material, defaulting to every placeholder the prompt renders except criteria. The vote's explanation is synthesised from the numbers that decided it (noul=0.92 (threshold 0.5)), and the full distribution lands on the judge span as judge.confidence and judge.probabilities. Settings a classify judge cannot use are named rather than dropped: system_prompt, prompt, temperature, structured_output, max_tokens, reasoning_effort, extra_kwargs and extra_body are listed in one WARNING when you set a non-default value, which says they still apply to prompted judges only when the panel actually seats one. Known classify ids warn at construction; models discovered through the fetched catalogue warn on their first call. repetitions > 1 warns because a classify verdict is deterministic; on a pairwise panel this means repetitions of each ordering, not the two distinct swapped orderings. Routing is decided per call by the Orq model catalogue's supports_classify flag, with a built-in list holding typesafe/jev-latest answering (and warning once) only when the catalogue has no entry at all. The built-in list and register_model() overrides provide synchronous hints for early validation and warnings, while the fetched catalogue remains authoritative at runtime, so a newly listed classify model works without manual registration. The endpoint lives on the Orq router, so a client that does not route through it cannot reach a classify judge. Red teaming and simulation are unchanged.
  • llm_config= on the simulation entry points — simulate(), generate_and_simulate(), generate(), generate_personas() / generate_scenarios() (and their singular forms), summarize_conversations(), datapoints_from_traces(), extend_from_traces(), extend_from_experiment() and wrap_simulation_agent() now take a full LLMCallConfig. It configures every simulation-side LLM call — five roles: the user simulator, the judge, the persona / scenario / first-message generators, the recommendations pass and the executive summary — and never the target under test, which is configured where it is constructed. max_tokens is narrower than the rest: the user simulator, the judge and the executive summary read it, while every generate_structured call site — the generators, the recommendations pass and the trace helpers — sizes its own budget from the item count and logs that the config's value was not used. llm_config.client is honoured on every one of those paths, so a caller who supplies their own client needs no credentials in the environment. Not exposed on the CLI: eq simulate still takes only --sim-model, which builds an LLMCallConfig naming that model, and the rest of the config is a Python-API surface. This is the seam that was missing once LLMCallConfig.temperature lost its default: a caller who wants to pin sampling on the judge (or set reasoning_effort, timeout_ms, extra_body) now has one object to pass instead of hand-building each agent. llm_config is the only place the simulation-side model is named — the sim_model= keyword that once shared that job is removed, see the entry under Notable defaults. Only the fields you set take effect, so an unset temperature still means the parameter is omitted from the request entirely. Mirrors red_team(llm_config=LLMConfig(attacker=..., evaluator=...)), which already had the equivalent surface. The user simulator and the judge take that config plus their prompt data as separate arguments, so a value you set explicitly to None — temperature=None, meaning "send no temperature" — reaches the request instead of collapsing back into "unset"; the deprecated AgentConfig and its subclasses are still accepted and still cannot express that distinction.
  • Token-usage tables in the red-team and simulation reports now report cached input tokens as a share of input (75 (75% of input tokens)) rather than a bare count. A raw cached figure says nothing about whether caching worked without the denominator it was cached against. Red-team reports gained the row outright — they carried only cache writes, which the Orq router reports as permanently 0, so the one measurable cache number was missing from the surface an operator reads. The row is omitted entirely when nothing was cached, so 0 (0% of input tokens) never reads as a measured result. This is a token share, not a cost share — a cached token bills at roughly a tenth of a fresh one and the provider exposes no cached-vs-uncached cost split, so the row says "of input tokens" and sits in the token table rather than beside Total Cost. Simulation's row was relabelled from Cached Tokens (retrieved) to Cached Input Tokens so both surfaces name the same number the same way.
  • extra_body= on the jury helpers — llm_jury(), llm_jury_pairwise() and PairwiseComparator now take extra_body= alongside extra_kwargs=. These helpers build their LLMCallConfig internally, so body fields previously had no seam on the jury path at all: extra_body is a structural field the call site owns, which means routing it through extra_kwargs raises — and since a judge failure becomes a verdict rather than propagating, the symptom was a judge that always failed. extra_kwargs still replaces a top-level call argument; extra_body is merged per key, so router-owned body fields survive alongside yours.
  • llm_jury() — LLM-as-a-jury evaluator for evaluatorq(evaluators=[...]). A single judge or a panel rates a target output against criteria; verdicts can be boolean (default), labeled categorical (labels= + passing_labels=), or numeric (verdict_kind="numeric" + threshold=). The panel consensus rule is selectable via aggregator=: "mode" (default) or "majority" (strict >50%) for categorical, "mean_std" (default) / "median" / "min" / "max" for numeric, or a custom Callable[[list[JuryVote]], ...]. Uses structured generation (tiered .parse → json_object fallback) and resolves the LLM client lazily on first scorer call so declaring an evaluator never requires credentials. The Responses-API path is deferred (RES-972). (RES-848)
  • OWASP_LLM_TOP_10 and OWASP_ASI_TOP_10 — public list[str] constants exported from evaluatorq.redteam. Pass them to red_team(categories=OWASP_LLM_TOP_10) to run a full framework sweep without spelling out individual category codes (RES-815).
  • simulate() and generate_and_simulate() accept a new opt-in upload_results= flag (default False). When set to True, results are uploaded to the Orq platform after the run, surfacing as an experiment when ORQ_API_KEY is configured. Upload errors are logged but never fail the call. Both functions also accept evaluation_description= and path= parameters mirroring evaluatorq() (RES-598).
  • LLMCallConfig — per-role LLM configuration with model, temperature, max_tokens, timeout_ms, extra_kwargs, and client fields
  • LLMConfig — now role-based via attacker: LLMCallConfig and evaluator: LLMCallConfig; retry, cleanup, and target-agent timeout settings retained at top level
  • LLMCallConfig exported from the evaluatorq.redteam public API
  • OpenAIModelTarget.send_prompt now enforces timeout_ms via asyncio.wait_for
  • Evaluator role config (temperature, max_tokens, timeout_ms, extra_kwargs, client) fully propagated through OWASPEvaluator, create_dynamic_evaluator, and create_owasp_evaluator
  • simulate() and generate_and_simulate() accept new evaluation_description= and path= parameters, forwarded straight to evaluatorq() (RES-598).
  • simulate() and generate_and_simulate() now run on top of evaluatorq(): persona × scenario datapoints are materialised, executed via a single evaluatorq job, and scored via adapted evaluators. This brings auto-upload, OTel tracing, the results table, CI gating, and dataset-id support to the simulation entry points "for free". The bespoke parallelism loop was removed; simulation/upload.py is kept as a standalone helper for direct callers but is no longer invoked from simulate() (RES-594).
  • simulate() accepts a new dataset_id= parameter — when set, simulation datapoints are streamed from the named Orq dataset (each row's inputs must already match a simulation input shape) instead of being passed inline. Mutually exclusive with datapoints and personas/scenarios (RES-594).
  • simulate() and generate_and_simulate() accept a new exit_on_failure= parameter, default True, for their own dropped-row gate. Evaluator score failures are returned in the results; dropped jobs raise SimulationDroppedError. Pass exit_on_failure=False for interactive / exploratory runs where you want dropped rows surfaced as warnings + error metadata instead of a non-zero exit (RES-594).

Bug Fixes

  • A call is now priced against the model that answered it rather than the id the caller asked for, falling back to the requested id when the served one is not in the catalogue, except for an orq/* router, whose call stays unpriced with a warning instead of billed at the router's rate. An orq/* router resolves to a different model per request, and the catalogue lists a single nominal rate for the router itself, so a run through one was billed at that headline rate no matter what served it. Measured on the live catalogue: orq/autorouter-anthropic-balanced resolved to eu.anthropic.claude-haiku-4-5 and was costed 4.55x over. When the served model cannot be priced, a router call now logs a warning that says why instead of reporting no cost silently. A prompt-cache breakpoint is now placed for an orq/* model too, for the same reason it is for agent/<key>: the resolved model is not knowable when the marker is written, and the router ignores rather than rejects the marker (RES-1573).
  • safe_substitute() dict keys were broken by Ruff RUF027 auto-fix in attack_generator, capability_classifier, and objective_generator — LLM prompts were receiving unsubstituted {placeholder} text, silently producing degraded attacks
  • generate_recommendations=True now correctly uses llm_config.evaluator.client before falling back to create_async_llm_client()
  • All hardcoded timeout literals (240_000, 90_000) replaced with config-driven values from LLMConfig / DEFAULT_TARGET_TIMEOUT_MS
  • OpenAITargetFactory now propagates max_tokens and timeout_ms to created targets
  • LangGraphTarget never populated ToolCallOutputItem.result: LangGraph returns tool results as separate ToolMessages and respond() skipped every message that was not an AIMessage. render_tool_call drops a call whose result is None, so judges saw transcripts with zero tool calls and any "did the agent use tool X" criterion failed regardless of what the agent did. Results are now paired with their call within the turn, as the pydantic-ai and OpenAI Agents targets already did. Two shapes still drop, both now warning: a call whose ToolMessage arrives in a later ainvoke (interrupt/resume graphs), and a call the graph emits without an id. A ToolMessage carrying Anthropic-style content blocks is unwrapped to its text rather than rendered as the block envelope, and a message dict keyed on type (BaseMessage.model_dump()) is now read rather than silently skipped.
  • A tool-call id from a non-OpenAI provider (toolu_* on Anthropic via LangGraph, run-* via pydantic-ai) was replayed as a Responses-API function_call.id, which rejects the whole request with "Expected an ID that begins with 'fc'" — killing the simulated user and the judge mid-run on any OpenAI-family model. Foreign ids are now dropped by the new responses_function_call_item_id() helper (exported from evaluatorq.openresponses.input_items). call_id is untouched and still carries the provider's id: it, not id, is what pairs a call with its output, and the API assigns an item id when none is sent.
  • new() on every target (LangGraphTarget, CallableTarget, CrewAITarget, VercelAISdkTarget, PydanticAITarget, OpenAIAgentTarget, OrqResponsesTarget, OpenAIModelTarget, ORQAgentTarget) constructs via type(self) instead of a hardcoded class name, so a subclass no longer silently degrades to the base class on every parallel job.

Internal

  • create_model_job split down to create_deployment_job: the router-model leg it also built was unreachable. parse_target returns only AGENT or DEPLOYMENT — llm:/openai:/direct: and unknown prefixes all raise — so _create_job_for_target's trailing create_model_job(model=value) could not execute, and the reasoning_effort forwarding inside that leg was dead with it. The fallback now raises for a kind with no leg rather than routing it to a job no target string could produce. Not exported from any __init__, so this is internal surface only.
  • usage_from_exception in common/structured_output.py replaces five hand-copied getattr(exc, 'usage', None) sites, each of which carried its own copy of the harvest rationale. A guardrail in tests/test_reuse_guardrails.py now fails on a sixth.
  • SaveMode converted from Literal to StrEnum
  • Timeout defaults centralised in contracts.py (DEFAULT_TARGET_TIMEOUT_MS = 240_000); PIPELINE_CONFIG import removed from openai.py and registry.py
  • MultiTurnOrchestrator.llm_kwargs constructor param deprecated — merged into _cfg.attacker.extra_kwargs at init time; use LLMCallConfig.extra_kwargs instead
  • RUF027 added to Ruff ignore list (intentional literal string keys used as safe_substitute template placeholders)
  • CLI --save flag migrated to typer.Choice
  • Ruff cleanup across all redteam modules (import sorting, Optional[X] → X | None, TYPE_CHECKING guards)

Breaking Changes (RES-877)

  • AgentTarget.send_prompt removed: respond(messages: list[Message]) -> AgentResponse is now the sole response method on every target; callers own the conversation transcript. Migrate target.send_prompt("x") to target.respond([Message(role="user", content="x")]).
  • OpenAIModelTarget, VercelAISdkTarget, and OpenAIAgentTarget are now stateless: per-instance _history is gone. Multi-turn conversation state is owned by the red-team orchestrator, not the target.
  • evaluatorq.redteam.ErrorInfo renamed to RunError: update any imports or isinstance checks that reference the old name.

Migration:

# Before
response = await target.send_prompt("Hello")

# After
from evaluatorq.contracts import Message
response = await target.respond([Message(role="user", content="Hello")])

New Features (RES-877)

  • AgentResponseError — a per-response error marker exposed on AgentResponse.error; used by the orchestrator to exclude failed turns from the replayed transcript.
  • turns_to_messages(turns, *, skip_errors=False) — helper exported from evaluatorq.redteam.contracts that converts a list of completed turns into a flat list[Message], optionally dropping turns whose response carries an AgentResponseError.
  • classify_error_type(error, *, existing_type=None) — exported from evaluatorq.redteam.contracts; infers a coarse error_type (content_filter, rate_limit, timeout, network_error, server_error, client_error, or unknown) from an error string. Shared by the orchestrator and report converters. On a per-response AgentResponseError, the orchestrator records an unmatched (unknown) result as target_error, so that field never carries unknown.
  • Tool-call fidelity on replay — the transcript replayed to a target now preserves assistant tool_calls and tool results across turns (OpenAIModelTarget as OpenAI chat params, VercelAISdkTarget as AI SDK CoreMessage tool-call/tool-result parts, OpenAIAgentTarget as Responses-API function_call/function_call_output items), so multi-turn tool-using agents see their prior tool context. VercelAISdkTarget accepts message_format="v5" (default) or "v4" to match the endpoint's AI SDK version (input/output:{type,value} vs args/result). Errored turns recorded by the orchestrator now carry a classified AgentResponseError.error_type instead of a flat target_error.

Internal (RES-899)

  • Unified tracing layer: the generic OTel span-recording helpers previously duplicated across redteam/tracing.py and simulation/tracing.py now live in a single evaluatorq.common.tracing module (truncate_for_span, capture_message_content, record_token_usage, record_llm_response, record_llm_input/output, set_span_attrs, get_trace_context_headers). Domain-specific span builders (with_redteam_span, with_simulation_span, with_llm_span) stay in their domain modules and import the shared helpers. The common module never imports from redteam, simulation, or openresponses.

Changed (RES-899)

  • Span PII gate env var renamed to EVALUATORQ_CAPTURE_MESSAGE_CONTENT (default true), replacing the previous OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT. The same name now gates both the Python and TypeScript simulation/red-team tracing layers. Set false / 0 to keep raw prompt and response text off spans (token usage, model, finish reason, and latency are still recorded).
  • Span text truncation defaults to capture-all in both Python and TypeScript. EVALUATORQ_SPAN_MAX_TEXT_CHARS is unset by default (no truncation); set a positive integer (canonical: 8192) to cap input and output message content, with the shared ... [truncated] marker. -1 / 0 / unset all mean capture all. The TypeScript path previously hardcoded a separate 2000-char cap with a … marker — both are gone.

Fixed (RES-899)

  • retry_statuses augments the default set again: passing a custom set (e.g. {429}) no longer silently drops the built-in 429 + 5xx retries — the custom statuses are added to the defaults, not substituted for them. (This restores the intended RES-897 review behavior, which was lost when #150 merged without the fix.)