Skip to content

Trace signals

A signal is a deterministic measurement or tag computed from an agent trajectory, with evidence showing which steps or tool calls contributed to it.

Use signals when you need repeatable measurements of an agent run's structure, tool use, or autonomy. They do not judge whether an answer is correct or helpful; use an evaluator or judge for answer quality.

Compute signals from a trace

compute_signals() takes an AtifTrajectory and returns a SignalReport. Convert a trace's raw Orq span records to ATIF first with OtelTrace.from_orq(...).to_atif(). The raw span records are the v3spans list from Orq's trace endpoint; evaluatorq.formats converts records but does not fetch them.

from pathlib import Path
from tempfile import TemporaryDirectory

from evaluatorq.formats import OtelTrace
from evaluatorq.signals import SignalsConfig, compute_signals

raw_spans = [
    {
        'span_id': 'chat-1',
        'parent_span_id': None,
        'name': 'chat-completion',
        'started_at': '2026-09-29T10:00:00Z',
        'ended_at': '2026-09-29T10:00:02Z',
        'attributes': {
            'gen_ai.operation.name': 'chat-completion',
            'gen_ai.response.model': 'eu.gpt-6.1-sol',
            'gen_ai.input.messages': [{'role': 'user', 'parts': [{'type': 'text', 'content': 'Hi'}]}],
            'gen_ai.output.messages': [{'role': 'assistant', 'parts': [{'type': 'text', 'content': 'Hello.'}]}],
        },
    }
]

trajectory = OtelTrace.from_orq(raw_spans).to_atif(agent_name='support-agent')

with TemporaryDirectory() as directory:
    config_path = Path(directory) / 'signals.json'
    config_path.write_text('{"tool_roles": {"FetchPage": "webview"}}', encoding='utf-8')
    config = SignalsConfig.from_file(config_path)
    report = compute_signals(trajectory, config=config)

print(report.values())
print(report.results['tool_error_count'].preconditions)
print(report.results['assistant_message_count'].evidence)

The example uses a local span fixture with the same shape as an Orq export, so it runs without credentials. To measure a recorded run, get its trace ID from the Orq trace URL, set ORQ_API_KEY, and fetch the span list:

TRACE_ID=trace_123
curl -fsS "${ORQ_BASE_URL:-https://my.orq.ai}/v2/traces/${TRACE_ID}/v3spans" -H "Authorization: Bearer ${ORQ_API_KEY}" > spans.json

Replace the fixture's raw_spans list with json.loads(Path('spans.json').read_text(encoding='utf-8')) after importing json. SignalsConfig.from_file() accepts a JSON object containing any subset of config fields; omitted fields keep their defaults, and unknown fields raise a validation error.

You can run one signal by passing only=['tool_error_count'] to compute_signals(). Group D tags load their metric dependencies automatically. report.values() returns only signals with a basis; inspect report.results[name] for evidence, preconditions, the approximation flag, or a no_basis explanation.

To use a signal as an evaluatorq scorer, create one with signal_evaluator('tool_error_count', threshold=2) or create all tag scorers with signal_evaluators(), then pass the result in evaluators=[...] to evaluatorq(). Group D tags fail when they fire; numeric measurements need a threshold to produce a pass or fail. signal_evaluators() also accepts names='all', one group with names='A' through 'D', or an explicit list of names.

from evaluatorq.signals import signal_evaluators

evaluators = signal_evaluators()  # Every trajectory tag, using the bundled thresholds.
print([evaluator['name'] for evaluator in evaluators])

Results and evidence

Each SignalResult carries a value, step and call evidence, preconditions, and an approximate flag. Evidence identifies the ATIF step and, for tool-call measurements, the call id; agent_path identifies the root or subagent trajectory. A failed precondition is recorded explicitly. If the signal cannot be computed, no_basis explains why and its value is None.

Some signals can be approximated when a trace lacks precise timestamps. For example, consecutive model-call step timestamps can bound an approximate model interval when OTel spans lack start and end pairs. The result keeps that value marked approximate=True. Compaction and copied-context steps are excluded from step and call measurements; embedded subagent trajectories still contribute to depth and invocation counts.

Inputs and coverage

Signals use information represented by the trajectory. Tool timing and tool definitions are preserved by the OTel-to-ATIF conversion when present in the OTel spans. Other ATIF sources can still support signals from their own fields, but they do not gain timing or schemas they never recorded. Missing tool result status, finish reasons, timestamps, or tool definitions can make a signal's precondition fail or qualify its value as approximate.

Jev classification is an optional preparation step for tools whose roles are not in SignalsConfig.tool_roles. classify_tool_roles(trajectories, config) classifies each unknown tool name once and returns a config with the roles merged in; it sends tool names to Jev without sending tool arguments. It uses SignalsConfig.classifier.model, which defaults to the classifier role (typesafe/jev-latest, set through EVALUATORQ_CLASSIFIER_MODEL or the signals task override; see Configuration › Models), and resolves a client from your configured credentials unless you pass client=; signal computation itself makes no model calls. Existing role entries take precedence, and unsuccessful classifications are recorded as other with a warning. You can instead provide role names directly in the JSON config.

SignalsConfig also controls error detection, argument canonicalisation, retry definitions, empty results, and tag percentiles. For example, the same JSON file can set {"error_detection": "status_and_content", "retry_definition": "same_tool_args_within_n", "retry_window": 3, "tag_percentiles": {"error_heavy.tool_error_rate": 90}} to detect errors in result text, look back three calls for retries, and use the cohort's 90th percentile for that tag clause. Its classifier.enabled field is descriptive configuration; call classify_tool_roles() explicitly when you want model-assisted classification.

Signal reference

The groups separate measurements by the kind of trajectory behavior they describe. Groups A–C return counts, ratios, durations, booleans, or per-tool breakdowns. Group D returns named tags whose fired rule clauses are recorded in reason and whose supporting steps and calls appear in evidence.

A — Structure

Signal Measures Additional input needed
max_depth Deepest subagent nesting level, with the root at depth zero. Embedded subagent trajectories
llm_call_count Explicit model-call counts, or one call for each agent step when no count is recorded. Agent steps; explicit call counts when available
user_message_count Number of user steps. None
assistant_message_count Number of agent steps. None
turn_count Number of user-to-agent turns. None
total_input_tokens Input tokens after subtracting cached tokens from prompt tokens. Token metrics
total_output_tokens Completion tokens across model calls. Token metrics
total_tokens Input plus output tokens. Token metrics
cache_read_token_share Cached tokens divided by prompt tokens. Cached and prompt token metrics
peak_context_tokens Largest prompt token count on one model call. Prompt token metrics
tool_call_count Number of tool calls. Tool calls
unique_tools_used Number of distinct tool names called. Tool calls
tool_call_value_count Number of calls for each tool name. Tool calls
bash_command_value_count Shell command calls by command family. Tool calls classified as bash
webview_value_count Web browsing and fetch calls. Tool calls classified as webview
loaded_skill_value_count Calls that load a named skill. Tool calls classified as skill
subagent_invocation_count Embedded subagent trajectories, including empty or unlinked children. Embedded subagent trajectories
total_subagent_messages Steps across embedded subagent trajectories. Embedded subagent trajectories
avg_messages_per_subagent_invocation Average subagent steps per invocation. Embedded subagent trajectories
model_count Number of distinct models used. Model name on the step or trajectory
provider_count Number of distinct providers inferred from model names. Model name on the step or trajectory
finish_reason_length_count Model calls that ended because of a length limit. Invocation finish reasons

B — Tools

Signal Measures Additional input needed
tool_error_count Tool results marked as errors. Tool results with status or error metadata
tool_error_rate Error results divided by tool calls. Tool results with status or error metadata
duplicate_tool_call_count Repeated calls with the same tool and canonical arguments. Tool calls
tool_retry_count Calls that repeat a prior failed call under the configured retry definition. Tool calls and error status
tool_succeeded_after_retry_count Retries that follow a failed call and then succeed. Tool calls and error status
invalid_schema_tool_call_count Calls that do not conform to a recorded tool definition. Tool calls and tool definitions
consecutive_same_tool_max Longest run of calls to one tool. Tool calls
consecutive_command_family_max Longest run of shell commands in one command family. Tool calls classified as bash
identical_tool_call_run_count Runs of identical tool calls. Tool calls
tool_oscillation_count Alternating tool-and-argument patterns, including two alternating argument sets on one tool. Tool calls and canonical arguments
tool_loop_count Detected repeated or oscillating tool-call loops. Tool calls
distinct_tool_arg_ratio Distinct canonical tool arguments divided by calls. Tool calls
empty_tool_result_count Results matching configured empty values or literals. Tool results
max_tool_result_bytes Largest tool result size in bytes. Tool results
total_tool_result_bytes Total tool result size in bytes. Tool results

C — Autonomy

Autonomous step counts include delegated subagent work between root user messages. max_autonomous_duration_ms uses root-agent timestamps, so a child step's later timestamp does not extend the root segment.

Signal Measures Additional input needed
autonomous_segment_count Runs of agent work between user messages. User and agent steps
max_autonomous_steps Most model responses plus tool calls in one autonomous segment. User and agent steps, tool calls
avg_autonomous_steps Average model responses plus tool calls per autonomous segment. User and agent steps, tool calls
max_llm_tool_cycles Longest consecutive run of tool-calling agent responses, measured separately along each agent path across the trace. Agent steps grouped by path, tool calls
terminal_answer_present Whether the final root agent step contains a terminal answer. Root agent steps
human_interruption_count User messages arriving before the root agent has given a terminal answer. User and agent steps
subagent_step_share Share of model responses plus tool calls occurring in subagent trajectories. Embedded subagent trajectories
subagent_message_share Share of messages occurring in subagent trajectories. Embedded subagent trajectories
parallel_tool_batch_count Agent steps containing multiple parallel tool calls. Tool calls grouped by step
max_parallel_tool_calls Largest parallel tool-call batch. Tool calls grouped by step
wall_time_ms Elapsed time across the recorded run. Start and end timestamps
active_time_ms Time spent in recorded model and tool operations. Model and tool timing
llm_time_ms Time spent in model calls. Model call bounds or consecutive model-call step timestamps
tool_time_ms Tool time that does not overlap model-call time. Tool result start and end timestamps
max_autonomous_duration_ms Longest root-agent segment duration between user messages. Root user and agent steps with timing

D — Tags

Tag Meaning
long_autonomous_run Extended execution without user input.
delegation_heavy Subagent use is high relative to the calibration cohort.
error_heavy Tool errors dominate execution.
tool_churn High tool activity coincides with low argument diversity or frequent retries.
tool_loop Repeated calls, oscillation, parameter drift, or a retry storm.
stalled The run ends without an answer after looping or repeated failure.
output_heavy Tool results are unusually large.
inefficient_execution At least two structural inefficiency indicators fire together.

The bundled thresholds are calibrated on local Claude Code sessions. The cohort field in tag_thresholds.json records the source, session counts, calibration date, and holdout count. These thresholds describe that cohort; treat the tags as cohort-relative indicators and recalibrate or override thresholds before applying them as universal limits to a different agent or workload.