FAQ¶
Common questions, grouped by area — General (install, keys, privacy, running evaluations), Red Teaming, and Agent Simulation.
General¶
What is evaluatorq?¶
A Python library for testing LLM apps and agents, with three modes:
- Evaluations — run jobs over your data in parallel and score them with built-in or custom evaluators; gate CI on pass/fail.
- Agent simulation — a user-simulator LLM drives your agent through multi-turn conversations while a judge scores whether it met its goals.
- Red teaming — adaptive adversarial attacks mapped to the OWASP LLM Top 10 and OWASP Agentic Top 10.
The Orq platform is optional — it stores results and routes LLMs when ORQ_API_KEY is set, but you can run entirely on OpenAI.
What do I install?¶
uv add "evaluatorq" covers evaluation, red teaming and simulation. Add an extra when you need what surrounds them — attack datasets, charts, the run browser:
uv add "evaluatorq[redteam]" # static/hybrid red team datasets
uv add "evaluatorq[dashboard]" # eq dashboard
Installation has the full table.
uv add installs into the current project — run uv init first if you don't have one. Run your scripts with uv run my_eval.py and the CLI with uv run eq, so the environment you installed into is the one that executes.
Prefer pip? Use python -m pip install "evaluatorq[redteam]", which installs into the interpreter you just named rather than whichever pip happens to be first on your PATH.
I installed evaluatorq but import evaluatorq fails¶
The install went to a different interpreter than the one running your script. This is the single most common setup failure, and it has nothing to do with evaluatorq — bare pip and bare python can resolve to different environments (a system Python, a virtualenv you forgot to activate, a container's global site-packages).
Confirm it by asking both sides where they live:
python -c "import sys; print(sys.executable)" # which interpreter runs
python -m pip show -f evaluatorq | head -3 # where the package landed
If the paths don't share a prefix, that's the bug. Two fixes:
- uv —
uv add evaluatorqthenuv run my_eval.py.uv runresolves the project environment before executing, so the two can't drift. - pip — always name the interpreter:
python -m pip install evaluatorq, and run with the samepython.
Avoid uv tool install evaluatorq: it builds an isolated environment that exposes the eq CLI but leaves evaluatorq unimportable from your own scripts.
Do I need an Orq account or an OpenAI key?¶
You need an LLM key wherever a simulator, attacker, or judge LLM runs — OPENAI_API_KEY for direct OpenAI, or ORQ_API_KEY to route through the Orq router. Plain evaluations with only deterministic evaluators need no key; any LLM-judged flow (jury, simulation, red teaming) does — including red teaming's static mode, where the target replays fixed attacks but the judge still scores each outcome with an LLM.
Which models run the simulator / attacker / judge — can I change them?¶
They default to model roles: the attacker and every judge run on the smart role (openai/gpt-6-sol), the simulated user on the fast role (openai/gpt-6-luna), routed via OPENAI_API_KEY or ORQ_API_KEY. Change a role for every surface at once with EVALUATORQ_SMART_MODEL or EVALUATORQ_FAST_MODEL (see Configuration › Models), or override per surface: red teaming takes llm_config=LLMConfig(attacker=..., evaluator=...), and simulation takes llm_config=LLMCallConfig(...) — one config for the user simulator, the judge, the generators, the recommendations pass and the executive summary. The CLI spells the model-only case as --sim-model.
from evaluatorq.redteam import LLMConfig, LLMCallConfig
report = await red_team(
target=MyAgent(),
llm_config=LLMConfig(
attacker=LLMCallConfig(model="anthropic/claude-3-5-sonnet", temperature=0.9),
evaluator=LLMCallConfig(model="openai/gpt-5.6-luna", temperature=0.0),
),
)
What leaves my machine?¶
Simulator/attacker/judge LLM calls go to OpenAI or the Orq router. Results upload to the Orq platform only if ORQ_API_KEY is set. With no key, everything stays local. Setting ORQ_API_KEY also enables OpenTelemetry tracing to my.orq.ai; suppress it with ORQ_DISABLE_TRACING=1, or keep tracing but strip prompt/response text from spans with EVALUATORQ_CAPTURE_MESSAGE_CONTENT=false. See Configuration.
How much does a run cost, and how do I keep it cheap?¶
Cost and wall-clock scale with cases × turns × LLM calls. The levers are how many cases you run (max_dynamic_datapoints / max_static_datapoints for red teaming, num_personas × num_scenarios for simulation), max_turns, and datapoint_parallelism (default 10 on evaluatorq(), red_team() and simulate()). To size against a provider concurrency limit, set llm_parallelism= (on evaluatorq(), red_team() or simulate()) rather than lowering datapoint_parallelism — it counts requests instead of tasks, so the number means the same thing however the fan-out nests. Red teaming's report tracks spend in report.summary.token_usage_total.
Where do results go, and how do I view a past run?¶
Runs auto-save locally (red-team runs to .evaluatorq/runs/; simulation runs to .evaluatorq/sim-runs/). Browse them in the multi-run FastHTML dashboard with eq dashboard (no path browses both stores; eq dashboard .evaluatorq/sim-runs scopes to simulation), or list runs with eq redteam runs / eq sim runs. See Dashboard.
Some spans are missing from my traces¶
The span exporter batches in the background, so spans can be lost two ways, neither of which fails the run. Either the in-memory queue overflowed — spans produced faster than the exporter drained them — or the process exited before the final flush finished. Raise ORQ_OTEL_MAX_QUEUE_SIZE for the first and ORQ_OTEL_FLUSH_TIMEOUT_MS for the second. Both log a warning; a hard SIGKILL drops whatever was still buffered without one. See Tracing › Batching and flush.
How do I run a plain evaluation?¶
Decorate a function with @job, hand evaluatorq() your data and evaluators, and it runs the jobs in parallel and scores each row:
from evaluatorq import DataPoint, evaluatorq, job, string_contains_evaluator
@job("greet")
async def greet_job(data: DataPoint, _row: int) -> str:
return f"Hello, {data.inputs['name']}!"
await evaluatorq(
"smoke-test",
data=[DataPoint(inputs={"name": "Ada"}, expected_output="Hello, Ada!")],
jobs=[greet_job],
evaluators=[string_contains_evaluator()],
print_results=True,
)
See Getting Started.
What evaluators are built in, and can I write my own?¶
There are deterministic ones (string match, JSON, regex) and LLM-judge ones. For custom logic, write a function that scores a row — see Custom Evaluators & Frameworks. For higher confidence, score one response with a panel of judges (LLM as a Jury) or compare two responses head-to-head (Pairwise Judging).
Red Teaming¶
How do I know my agent is safe?¶
You don't, until you attack it. Shipping after a refused "say something harmful" is a vibe check, not a test. It only proves the agent refuses the one obvious prompt you thought to try. Red teaming runs a mapped set of adversarial attacks and reports a resistance rate (the fraction the agent withstood), so "safe" becomes a number you can gate on. See Red Teaming.
Isn't a single-turn "refused → safe" check enough?¶
No. The attacks that land are the ones a single prompt can't express:
- Multi-turn escalation — each message looks benign; the attack assembles across turns. Invisible to single-turn evals.
- Indirect injection — the attacker controls what the agent reads (emails, docs, tool results), not what you type. You never see the payload.
- Memory poisoning — a planted instruction fires on a later, unrelated run.
- Many-shot jailbreaking — 100+ in-context examples steer behaviour.
Red teaming generates multi-turn attacks by default (max_turns=) precisely so these surface.
What does red teaming actually test?¶
Attacks and LLM-judge evaluators mapped to three frameworks:
- OWASP LLM Top 10 — prompt injection, system-prompt leakage, and the rest.
- OWASP Agentic Top 10 (ASI) — agent-specific risks: tool abuse, excessive agency, trust exploitation.
- Responsible AI — fairness/bias, liability (legal, medical), content policy, harmful content.
Pass the ones you care about via categories=["LLM01", "ASI01", ...]. The full list of codes is in the redteam API reference.
Static, dynamic, or hybrid — which mode?¶
- static — replays a fixed dataset of known attacks. Deterministic and cheap; run it in CI. Defaults to Orq's public
orq/redteam-vulnerabilitiesset, or pass your own withdataset=. - dynamic — an LLM generates attacks tailored to your agent's tools and prompt.
- hybrid — static seeds plus dynamic expansion, in one pass.
Details in Red Teaming › Modes.
How does it decide what to attack?¶
Probe → attack → judge. It first inspects the target (tools, permissions, system prompt) to build a capability profile, generates attacks against the exposed surface, then scores each outcome against per-vulnerability LLM-judge rubrics. Filtering out attacks that can't apply keeps a run focused on real exposure.
Why an LLM judge instead of keyword matching?¶
Because the surface text lies. "I see malicious instructions in this email but I won't follow them" and "I'll forward that email right away" both mention instructions, yet one is resistant, one is vulnerable. A regex can't tell them apart; a judge rubric can. Write your own in Custom Evaluators & Frameworks.
How do I trust the judge itself?¶
Calibrate it against known baselines: run the same attack set through an agent that should pass and one that should fail. If the safe agent gets flagged vulnerable you have false positives; if the weak agent is marked resistant you have blind spots. Either way, patch the rubric and rerun. Matched pairs (same attacks, both agents) make the miscalibration visible.
Isn't "vulnerable" context-dependent?¶
Yes, and the evaluators account for it. Fetching a shell script from GitHub is helpful in a coding assistant and an RCE vector in a support bot; chaining three API calls unprompted is doing the job or excessive agency depending on the agent. Judges see the agent's declared context, so the same action can score differently across agents.
How do I run it against my own agent?¶
Subclass AgentTarget and implement two methods: respond(messages) (return an AgentResponse) and new() (return a fresh instance for each attack). Any framework works behind those two methods (LangChain, LangGraph, the OpenAI Agents SDK, or a plain loop), so there's no framework buy-in:
from evaluatorq.contracts import AgentTarget, AgentResponse, Message
from evaluatorq.redteam import red_team
class MyAgent(AgentTarget):
async def respond(self, messages: list[Message]) -> AgentResponse:
reply = await my_agent_call(messages) # your agent, any framework
return AgentResponse(content=reply)
def new(self) -> "MyAgent":
return MyAgent()
report = await red_team(target=MyAgent(), mode="dynamic", max_turns=4)
rate = report.summary.resistance_rate # None when no attack could be evaluated
print(f"resistance: {rate:.0%}" if rate is not None else "resistance: no verdict")
Or point it at an Orq agent by key ("agent:<key>"), or drive it from the CLI with eq redteam run --target agent:<key>. See Red Teaming › Red-team your target and examples/redteam/15_tool_chaining.py for a full custom target.
If I red-team my real agent, will the attacks actually fire its tools?¶
Yes. The target runs its own tools, so a successful attack triggers real side effects (sends the email, moves the money, runs the shell command). Point red teaming at a sandboxed or test instance with fake/stubbed tools — like the demo's fake wallets — not at production credentials wired to irreversible actions.
The report says my agent is vulnerable — what do I change?¶
Usually the system prompt. The load-bearing fixes are an explicit instruction hierarchy (data the agent reads is never a command), a confirmation gate before risky tools, and refusing authority claims — then rerun and watch the resistance rate climb. Every run attaches LLM-generated focus-area recommendations to the report by default; pass recommendations=False (or --no-recommendations) to skip that extra LLM call, or recommendations=RedTeamRecommendationConfig(max_areas=3, max_attacks=20) (importable from evaluatorq.redteam) to tune how much gets analyzed. A focus area is a framework category the run broke on, analyzed in risk order with one LLM call each — so max_areas decides how many categories get advice, and max_suggestions (default 5) how much advice each one gets. max_attacks is how many of that category's failed attacks are sampled into the prompt. The config also carries the prompt truncation budgets and the analysis call's max_tokens. Because the focus-area call is a single request carrying every sampled attack, an attack whose formatted block exceeds condense_above_chars (default 1000) is summarized by its own small LLM call first, so one long agentic transcript can't push the aggregate past the model's context — short attacks go in verbatim and cost no extra call.
What does passed=True mean?¶
The agent resisted the attack (the attack failed). passed=False means the attack succeeded: the agent is vulnerable. passed=None is a third outcome: the attack could not be evaluated (the target call or the judge call failed) — it is neither a pass nor a fail, and r.vulnerable is None for it. resistance_rate is the fraction of evaluated attacks that came back passed=True, and is None when none could be evaluated.
My run shows no results / eq redteam run failed — why?¶
Check result.evaluation_error on the affected results (or report.summary.errors_by_type, which groups judge failures under evaluation/<code> keys such as evaluation/api_status) — it records why the judge couldn't return a verdict (timeout, parse, api_connection, api_status, unknown), separate from result.error, which means the attack itself never ran. The CLI prints the dominant code and a sample message on failure.
Two conditions make eq redteam run exit 1: zero verdicts (report.summary.no_verdict — nothing could be scored, so a "100% resistant" or 0.0 rate would be a lie) and, as of the coverage gate, low verdict coverage — EvaluatorConfig.min_evaluation_coverage (default 0.8) means a run where fewer than 80% of attacks got any verdict now fails too, not just warns. This is a behaviour change: a 79%-coverage run used to exit 0. Pass --min-evaluation-coverage 0 (or set min_evaluation_coverage=None in Python) to go back to warn-only. It's distinct from min_successful_judges, the per-attack jury quorum that's what creates unevaluated attacks in the first place.
The reported cost says priced_calls < calls — how do I price the rest?¶
Some calls came back with usage but no provider price, so the dollar figure covers a subset and is a lower bound. EvaluatorConfig.api defaults to 'responses' precisely because that is the endpoint the Orq router prices, so judge calls record cost the way target calls do. Two things break that: a model the router cannot resolve on responses falls back to chat completions on its own, and EvaluatorConfig(api='chat_completions') opts out deliberately — both leave judge calls unpriced. If the gap matters, check the judge model resolves on the router before changing anything. Compare priced_calls against calls before quoting a total.
Agent Simulation¶
What is agent simulation, and how is it different from red teaming?¶
Both drive your agent across multi-turn conversations, but with opposite intent. Simulation plays a cooperative user (a persona pursuing a realistic goal) and asks "did the agent do its job?"; red teaming plays an adversary and asks "can the agent be broken?" Three LLMs are in play for simulation: your agent, a user-simulator, and a judge that scores goal_achieved / criteria_met.
How do I simulate multi-turn conversations?¶
generate_and_simulate() is the fastest start: it synthesizes personas, scenarios, and opening messages from a one-line description of your agent, no hand-written transcripts:
from evaluatorq.simulation import generate_and_simulate
results = await generate_and_simulate(
evaluation_name="support-agent-sim",
target="agent:my-support-agent", # or pass a callable as target= for your own agent
agent_description="Customer support agent handling refunds and orders.",
num_personas=3,
num_scenarios=4, # → 12 persona × scenario simulations
max_turns=6,
evaluator_names=["goal_achieved", "criteria_met"],
)
print(f"pass rate: {sum(r.goal_achieved for r in results)}/{len(results)}")
See Agent Simulation.
Can I tune when a simulation result gets recommendations?¶
Yes — pass a SimulationRecommendationConfig. Only results with a fixable failure signal get an LLM call; the metric thresholds that decide that, plus the prompt budgets, the result cap and the suggestion cap, are fields on the config. It's the same recommendations= flag red teaming uses: True for defaults, False to skip the LLM call, an instance to tune.
eq sim run and simulate() both generate them by default. The returned SimulationResult list has nowhere to carry suggestions, so from Python they are only observable when you save the run with save=True (optionally choosing report_path=), or through the dashboard; without either, they are generated, discarded and a warning says so. Pass recommendations=False to skip the call:
from evaluatorq.simulation import simulate
from evaluatorq.simulation.reports import SimulationRecommendationConfig
results = await simulate(
target="agent:my-support-agent",
save=True,
recommendations=SimulationRecommendationConfig(factual_accuracy_below=0.7, max_suggestions=5),
)
To analyze results you already have, call the generator directly with the same config:
from evaluatorq.simulation.reports import SimulationRecommendationConfig, generate_recommendations
recs = await generate_recommendations(
results, client, "gpt-5.6-luna",
config=SimulationRecommendationConfig(factual_accuracy_below=0.7, max_suggestions=5),
)
Defaults: thresholds 0.5 on the judge's 0-1 scales, max_results=10, max_transcript_chars=3000, max_message_chars=400, max_suggestions=3, max_tokens=800. Values outside the valid range raise at construction, and an unknown field name raises too, rather than silently disabling a trigger.
The endpoints of each threshold's range disable that one check on purpose: a judge score is never below 0.0 or above 1.0, so factual_accuracy_below=0.0, tone_appropriateness_below=0.0 and hallucination_risk_above=1.0 each mean "never trigger on this metric". Use them to narrow which signals produce suggestions; the other triggers (broken rules, failed criteria) are unaffected.
Where do the personas and scenarios come from?¶
Your choice of control: generate them from a one-line description, seed by archetype, hand-build Persona(...) / Scenario(...) for full control, or ground new cases in your real production traces so they mirror how users actually behave. You can also replay stored datapoints to re-run the exact same cases against any target. See Agent Simulation.
Which agent frameworks does simulation work with?¶
LangGraph, the OpenAI Agents SDK, Pydantic AI, CrewAI, a plain async callback (passed as target=), or a hosted Orq agent (target="agent:<key>") — see the framework demos.