Configuration Reference¶
Complete reference for every configuration surface in orq-arena: the environment variables
it reads, the YAML config files under the project root and configs/, and the prompts
file format. The tool is configured by two layers:
- A
.envfile holding the single secret orq-arena needs (ORQ_API_KEY), loaded at CLI startup. - A YAML file (
orq_arena.yamlby default) describing match rules, the gateway client, the candidate pool, and the judge panel, parsed and validated into anArenaConfigPydantic model.
The source of truth for every key below is src/orq_arena/config.py (models MatchRules,
OrqAIGatewayConfig, PreflightConfig, ArenaConfig) and src/orq_arena/candidates.py
(CandidateSpec). All file paths in this document are relative to the project root.
Environment Variables¶
Copy the template and fill it in before running anything that calls a real model:
| Variable | Required | Default | Description |
|---|---|---|---|
ORQ_API_KEY |
Required for live runs | (none) | The only secret orq-arena needs. Every candidate, judge, and preflight-probe call goes through the orq.ai router gateway with this one key; the run fails up front with ORQ_API_KEY is not set. Export it before running orq-arena. if it is missing. Create one in your workspace settings, as .env.example points to, or per the API keys guide. |
ORQ_BASE_URL |
No | https://api.orq.ai |
Points completions and the model/price catalog at a different orq.ai host (staging, a proxy). Honoured only while gateway.base_url is left at its default: setting that key in the YAML is a bring-your-own-endpoint opt-out that wins outright, so the run can never be split between two hosts. Not read from .env.example; export it yourself when you need it. |
Notes:
ORQ_API_KEYis not required fororq-arena pool(prints the candidate pool, never constructs a gateway) or the log-reading commands (report,annotate,anchor). Note thatreportdoes use it when present, for the one catalog read that prices the cost section; without a key the page renders with that section omitted.
.env loading¶
Every CLI invocation reads ./.env once before any subcommand runs. It parses KEY=VALUE
lines (skipping blanks and # comments, stripping surrounding quotes) and loads them with
os.environ.setdefault, so the real shell environment always wins; .env only fills in
what the shell hasn't set. A missing .env is silently fine.
.env is git-ignored; .env.example is the committed template with ORQ_API_KEY= blank.
Config Files¶
| File | Purpose |
|---|---|
orq_arena.yaml |
The default model pool + rules, shipped at the project root. Loaded whenever --config is omitted or points here, DEFAULT_CONFIG = "orq_arena.yaml" in src/orq_arena/cli.py. Ships 8 candidates, uniform thinking-OFF, so the ELO compares models rather than vendor default reasoning settings. |
configs/reasoning_arena.yaml |
Uniform thinking-ON counterpart of the default file, the "does thinking help?" benchmark. Not loaded automatically; run it explicitly with --config configs/reasoning_arena.yaml. |
configs/frontier_8.yaml, configs/budget_8.yaml, configs/frontier_16.yaml |
Ready-made pools tiered by the Artificial Analysis intelligence index (frontier ~40-56, budget ~12-25, 16-model stress test). See configs/README.md. |
Any YAML path can be passed to --config; load_config() (src/orq_arena/config.py) reads it
with yaml.safe_load and validates it into an ArenaConfig via
ArenaConfig.model_validate(raw). There is no ${VAR} environment-variable substitution inside
the YAML, every value is literal. Unknown keys are rejected, not ignored: a typo such as
replacment_judges: fails the load with a did-you-mean suggestion instead of silently
configuring nothing.
Which commands take --config¶
| Command | --config behavior |
|---|---|
orq-arena run |
Required, no default. The YAML candidates are used as-is; the run is headless by default (--tui opts into the live show). |
orq-arena pool |
Required, no default. Prints the configured candidate pool. |
orq-arena rejudge |
Defaults to orq_arena.yaml. Supplies gateway and (unless --criteria overrides it) criteria. |
orq-arena report |
Reads run identity from the log's *.run.json manifest. orq_arena.yaml is only a fallback for a log with no manifest; passing --config explicitly overrides the manifest and is disclosed on the page. |
orq-arena refresh-catalog |
Defaults to orq_arena.yaml. Only gateway is used, to re-fetch the workspace model catalog. |
orq_arena.yaml field reference¶
A minimal example, drawn from the real shipped file (some candidates entries and all-default
sections omitted for brevity, the code default is noted inline where a key isn't set in the
shipped file):
match:
max_rounds: 5
# starting_hp / damage_unanimous / damage_majority are TUI-presentation-only
# knobs (the rating never sees them); they take code defaults if omitted.
starting_hp: 100
damage_unanimous: 30
damage_majority: 15
gateway:
base_url: https://api.orq.ai/v3/router
candidate_max_tokens: 2048
judge_max_tokens: 2048
stream_read_timeout_s: 1200
judge_timeout_ms: 90000
# preflight and headless_concurrency not set in the shipped
# file - both take their code defaults (see tables below).
candidates:
- model_id: anthropic/claude-opus-4-8
- model_id: openai/gpt-5.4
- model_id: google/gemini-3.1-pro-preview
reasoning: { thinking: { type: disabled } }
# ...5 more candidates in the shipped orq_arena.yaml
judges:
- anthropic/claude-haiku-4-5-20251001
- google/gemini-2.5-flash-lite
- openai/gpt-5.4-nano
replacement_judges:
- mistral/mistral-medium-2604
criteria: >-
Accuracy and correctness, helpfulness and completeness, clarity, and
relevance to the prompt.
min_successful_judges: 2
match (MatchRules)¶
Only max_rounds affects what gets rated. starting_hp, damage_unanimous, and
damage_majority are TUI-presentation-only knobs. Hit points are the fighting-game
style health bar each model carries in the live show; the engine no longer tracks them, so
these keys feed nothing but that display. The TUI recomputes hit points, damage tiers, and
knockouts client-side from the judged verdicts; the rating never sees any of them.
| Key | Type | Default | Effect |
|---|---|---|---|
max_rounds |
int |
5 |
Prompt cap per match. The actual number of rounds run is min(max_rounds, len(prompts)) (preflight.call_counts), a smaller prompts file also shortens matches. The only match key that affects the rating. |
starting_hp |
int |
100 |
TUI-only. Hit points each candidate's health bar starts a match with, in the live show. Never read by the rating. |
damage_unanimous |
int |
30 |
TUI-only. Hit points the show's bar subtracts when a round's decisive votes agree unanimously. Presentation drama only; the rating is decided by judged round wins, not hit points. |
damage_majority |
int |
15 |
TUI-only. Hit points the show's bar subtracts on a decisive but non-unanimous (split) verdict. Presentation drama only. |
Verdict pacing is no longer a config key: the seconds the TUI holds on each verdict before the
next round is a code constant, VERDICT_HOLD_S in src/orq_arena/tui/hp.py (headless runs never
pause). The match winner is the side with more judged round wins (equal round wins is a draw, an
empty winner), and every prompt in the drawn slice is always judged regardless of where the
on-screen health bar happens to sit.
gateway (OrqAIGatewayConfig)¶
| Key | Type | Default | Effect |
|---|---|---|---|
base_url |
str |
"https://api.orq.ai/v3/router" |
Base URL for the AsyncOpenAI client (OrqGateway.__init__, src/orq_arena/providers/orq_gateway.py). One OpenAI-compatible endpoint fronts every provider, models, judges, and the preflight probe all share it. Left at the default, host resolution is delegated to evaluatorq and honours ORQ_BASE_URL; changing it here opts out of that entirely (see Alternate configs and overrides). |
candidate_max_tokens |
int |
2048 |
Default per-response output cap for candidate completions (stream_completion's max_tokens=max_tokens or self._cfg.candidate_max_tokens). Too low truncates long or creative answers, a cut response is flagged โ truncated in the TUI response panel (src/orq_arena/tui/widgets/response_panel.py) and judges tend to penalize it. Overridden per-candidate by candidates[].max_tokens. |
judge_max_tokens |
int |
2048 |
Output cap for judge calls, passed to evaluatorq's llm_jury_pairwise(max_tokens=...). A cap, not a target: it costs nothing extra on frugal judges. Thinking-by-default judges (e.g. gemini-2.5-flash) burn reasoning tokens before writing a verdict; a low cap starves the verdict entirely and fails the vote (the codebase's own regression case: 512 produced a LengthFinishReasonError on every one of that judge's votes). 2048 leaves headroom without materially raising cost on the cheap default panel. |
stream_read_timeout_s |
int |
1200 |
Max silence between stream chunks, in seconds, before the client treats the connection as dead (httpx.Timeout(read=float(stream_read_timeout_s), ...) in OrqGateway.__init__). This is a read-gap timeout, not a total-duration cap, a thinking model that pauses for minutes before its first token is fine as long as chunks keep arriving within this gap. A stream that goes silent longer than this is retried once, then the round is voided (logged, shown, excluded from scoring). 1200s = 20 minutes, deliberately generous. |
judge_timeout_ms |
int |
90000 |
Per-judge-call timeout in milliseconds, passed to evaluatorq's llm_jury_pairwise(timeout_ms=...) (Battle.__init__, src/orq_arena/arena/battle.py). 90000 (90s) is also evaluatorq's own library default for this parameter. |
Not configurable via YAML: connect=10.0, write=60.0, and pool=60.0 second timeouts are
hardcoded in OrqGateway.__init__ alongside stream_read_timeout_s, only the read-gap timeout
is exposed as a config key. The preflight probe call (see below) also uses a hardcoded
max_tokens=1000, independent of candidate_max_tokens/judge_max_tokens.
preflight (PreflightConfig)¶
| Key | Type | Default | Effect |
|---|---|---|---|
thinking_probe |
bool |
True |
Before a live run, sends one tiny probe call ("Reply with the single word: ok") per candidate to detect vendor-default thinking that contradicts its reasoning config (thinking_probe, src/orq_arena/preflight.py). Surfaces as ๐ง โฆ thinks despite config in the CLI preflight output and footnotes the leaderboard for that candidate. Adds one extra call per candidate (probe_calls in preflight.CallCounts). Set false to skip those extra calls. |
Top-level run settings (ArenaConfig)¶
| Key | Type | Default | Effect |
|---|---|---|---|
headless_concurrency |
int |
4 |
Matches run in parallel under an asyncio.Semaphore(max(1, headless_concurrency)) on headless runs, the default (run_headless โ run_tournament(concurrency=...), src/orq_arena/headless.py). The TUI (--tui) always passes concurrency=1 internally so the live show stays one fight at a time. |
candidates (the model pool)¶
candidates: list[CandidateSpec]: required at the top level, and the parsed list must contain at
least 2 entries (ArenaConfig._validate: "Need at least 2 candidates, got {n}").
| Key | Type | Default | Effect |
|---|---|---|---|
model_id |
str |
(required) | orq.ai router gateway model slug, e.g. anthropic/claude-opus-4-8. The only required field per candidate entry. |
name |
str |
"" โ falls back to short_model |
Display name used on the leaderboard, TUI cards, and arena events (MatchStarted/MatchResolved). Defaults to model_id with the provider prefix stripped ("anthropic/claude-opus-4-8" โ "claude-opus-4-8") and is generated only to break a collision: when two candidates would answer to the same name, both fall back to their full model ids and the preflight says so. A custom name is honoured as written. Note: battles.jsonl records store both short_model (model_a/model_b) and the full router id (model_a_id/model_b_id); the id is what every per-model view keys on, because short names collide across providers. name is presentation-only. |
emblem |
str |
"" |
Optional glyph/emoji shown before the orc name on the TUI candidate card (src/orq_arena/tui/widgets/model_card.py). Purely cosmetic. |
reasoning |
dict \| null |
None |
Raw router reasoning-control object, forwarded verbatim as extra_body on the completion request (stream_completion, src/orq_arena/providers/orq_gateway.py). Not interpreted beyond the budget_tokens cross-check below, the router normalizes it per provider. |
max_tokens |
int \| null |
None โ falls back to gateway.candidate_max_tokens |
Per-candidate override of the response output cap. |
Reasoning passthrough recipes (forwarded untouched; the router normalizes per provider,
comment block in orq_arena.yaml):
# OpenAI -> reasoning: { reasoning_effort: low|medium|high }
# Claude -> reasoning: { thinking: { type: enabled, budget_tokens: 4096 } }
# Gemini 3 -> reasoning: { thinking: { thinking_level: low|high } }
To explicitly disable a think-by-default model (e.g. some Gemini models) for a uniform thinking-OFF pool:
Cross-field validation: if reasoning.thinking.budget_tokens is set, it must be strictly less
than the candidate's effective cap (max_tokens if set, else gateway.candidate_max_tokens), a
config where the thinking budget meets or exceeds the output cap fails to load with
ValueError: {name}: thinking budget_tokens ({budget}) must be < max_tokens ({cap})
(ArenaConfig._validate, src/orq_arena/config.py).
The shipped orq_arena.yaml also documents which models it deliberately excludes from the
default pool because the router can't disable their thinking (see the comment below the candidates list
in that file), those belong in configs/reasoning_arena.yaml instead.
judges, replacement_judges, criteria, min_successful_judges¶
| Key | Type | Default | Effect |
|---|---|---|---|
judges |
list[str] |
(required, non-empty) | Router model ids forming the base judge panel handed to evaluatorq's llm_jury_pairwise(judges=...). Must be non-empty, ArenaConfig._validate raises ValueError("Judge panel is empty") otherwise. Each judge votes on both seat orderings of every round. |
replacement_judges |
list[str] |
[] |
Neutral stand-ins promoted when a primary judge errors mid-run (llm_jury_pairwise(replacement_judges=...)). |
criteria |
str |
"Accuracy and correctness, helpfulness and completeness, clarity, and relevance to the prompt." |
Free-text criteria string every judge is given, what the jury is asked to judge on. Can be overridden for a single re-judge without editing the YAML via orq-arena rejudge ... --criteria "...". |
min_successful_judges |
int |
2 |
Minimum number of decisive reconciled votes required for a round to produce a real verdict. Fewer than this, and the round is inconclusive (dropped from the rating, doesn't count toward max_rounds), a guard against a "jury of one" deciding a round. |
Self-judge exclusion: per match, any judge whose model_id matches either contestant's
model_id is filtered out of that match's panel (panel = [m for m in cfg.judges if m not in
contestants], Battle.__init__). If that empties the panel entirely, Battle.__init__ raises
ValueError("Every judge is a contestant in ..., add a neutral judge to the config."): a
small judges list can strand a specific pairing if both models in that match are also
configured as judges elsewhere in the config.
Prompts file format¶
Prompts are not part of ArenaConfig, the file is a separate --prompts CLI flag on
orq-arena run (default prompts/starter.jsonl, DEFAULT_PROMPTS in src/orq_arena/cli.py). Format: one
JSON object per line (JSONL), parsed by load_prompts() into a list of PromptItem
(src/orq_arena/data/prompts.py).
Instead of a file path, --prompts orq:<dataset_id> pulls the datapoints of an
orq.ai Dataset through the orq-python
SDK, authenticated with the same ORQ_API_KEY as the gateway.
Mapping per datapoint: the last user message becomes the prompt text, {{var}} placeholders
are filled from the datapoint's inputs, multi-part content is joined, and datapoints without
a user message are skipped (the count is reported if the dataset yields nothing usable).
A datapoint's inputs.category becomes the prompt's category, falling back to general when
the datapoint doesn't set one; expected_output is not read today.
Each prompt carries its source datapoint_id in prompt_metadata, so every round in
battles.jsonl joins back to the exact datapoint that produced it.
When --prompts orq:<dataset_id> is used, the run also captures the dataset's identity:
orq_dataset_meta() (src/orq_arena/data/prompts.py) returns {id, name, url}, an orq.ai
studio link plus a display name fetched via the SDK's datasets.retrieve on a best-effort
basis (any failure, offline included, leaves the id standing in as the name, so a run is
never blocked on this call). This lands under a dataset key in battles.run.json, present
only for dataset-sourced runs, and battles.report.html links the dataset by that name.
| Field | Type | Required | Effect |
|---|---|---|---|
prompt |
str |
Required (or text, see below) |
The prompt text, becomes PromptItem.text. |
text |
str |
Fallback for prompt |
Read only if prompt is absent (row.get("prompt") or row.get("text")). A row with neither key is silently skipped. |
category |
str |
Optional, default "general" |
Feeds the per-category ELO slices on the leaderboard. Untagged rows land in "general". |
| any other key | any | Optional | Carried verbatim as prompt_metadata on every battle record for that prompt in battles.jsonl, an opaque pass-through for joining rounds back to your source data (never sliced or judged on). |
PromptItem itself (src/orq_arena/data/prompts.py) is a frozen dataclass with three
fields: text: str, category: str = "general", and metadata: dict = {}.
The shipped prompts/starter.jsonl has 30 prompts across four categories: code (8), general
(11), math (6), and creative (5). Example row:
{"prompt": "Write a Python function that finds the longest palindromic substring in a given string. Explain your approach.", "category": "code", "length_bucket": "medium"}
Note: every row in the shipped file carries a length_bucket key (short/medium), it
is not read by load_prompts() today and has no effect on the run.
Required vs Optional Settings¶
Settings that cause a hard failure (config load or first live call) if absent or invalid:
candidates: required top-level key, must parse to at least 2CandidateSpecentries, and each entry requiresmodel_id. Missing/short lists failArenaConfigvalidation immediately.judges: required top-level key, must be a non-empty list.- Any
candidates[].reasoning.thinking.budget_tokensmust be strictly less than that candidate's effectivemax_tokens(own override orgateway.candidate_max_tokens), or config loading fails. ORQ_API_KEYmust be set in the real environment, not required to load the YAML, but the gateway raisesRuntimeErrorthe moment any live call is attempted (orq-arena run,rejudge,refresh-catalog). Not needed fororq-arena pool,report,annotate, oranchor.- At least one configured judge must not be a contestant in a given match, or that match raises
ValueErrorat battle start.
Everything else is a Pydantic default and safe to omit from the YAML entirely:
| Key | Default |
|---|---|
match.max_rounds |
5 |
match.starting_hp (TUI-only) |
100 |
match.damage_unanimous (TUI-only) |
30 |
match.damage_majority (TUI-only) |
15 |
gateway.base_url |
https://api.orq.ai/v3/router |
gateway.candidate_max_tokens |
2048 |
gateway.judge_max_tokens |
2048 |
gateway.stream_read_timeout_s |
1200 |
gateway.judge_timeout_ms |
90000 |
preflight.thinking_probe |
true |
headless_concurrency |
4 |
replacement_judges |
[] |
criteria |
"Accuracy and correctness, helpfulness and completeness, clarity, and relevance to the prompt." |
min_successful_judges |
2 |
candidates[].name |
short model id, or the full id when two candidates would share one |
candidates[].emblem |
"" |
candidates[].reasoning |
null |
candidates[].max_tokens |
null (โ gateway.candidate_max_tokens) |
Alternate configs and overrides¶
orq-arena is a CLI tool, not a deployed service, so it has no dev/staging/production split of its own. The axes for changing behavior between runs are:
- Different model pool / benchmark question: pass a different YAML to
--config. The two shipped presets areorq_arena.yaml(uniform thinking-OFF, the default) andconfigs/reasoning_arena.yaml(uniform thinking-ON), run the latter withorq-arena run --config configs/reasoning_arena.yaml.load_config()accepts any path.orq-arena refresh-catalog --showlists your workspace-enabled model ids to paste into thecandidateslist. - Different jury on an already-recorded run, no regeneration:
orq-arena rejudge <log_path> --judge <id> [--judge <id> ...] [--criteria "..."]re-scores the responses already inbattles.jsonlwith a new panel and/or criteria, without touching the YAML file. - Different gateway host: two ways, and they do not mix.
ORQ_BASE_URLin the environment retargets an otherwise-default run: completions go through evaluatorq's shared resolver, and the catalog/price read follows the same host (catalog_host,src/orq_arena/providers/models_list.py). This is the staging path.gateway.base_urlin the YAML is the bring-your-own-endpoint opt-out. Setting it to anything other than the default makes the YAML win andORQ_BASE_URLis ignored, for both completions and the catalog, so a run is never priced against one host while calling another.
Regenerated / git-ignored files¶
These are run outputs, not configuration, do not hand-edit them as config, and note they are
git-ignored (.gitignore):
| File | Written by |
|---|---|
.env |
Hand-authored from .env.example; never committed. |
battles.jsonl |
orq-arena run, one row per judged round (BattleRecord, schema v4, written as each round resolves; includes per-model ttft_a_ms/ttft_b_ms and duration_a_ms/duration_b_ms timing fields). |
battles.run.json |
orq-arena run, the run manifest (the config it ran, config + prompt hashes, the prompts path, the host the run actually used, panel, seed, agreement stats; also a dataset key with id, name, and studio URL for dataset-sourced runs). Holds no credential: ORQ_API_KEY is read from the environment and never enters the config. |