Formats¶
evaluatorq.formats converts one agent run between four shapes: a chat message list, OpenResponses items, an OpenTelemetry GenAI trace and an ATIF trajectory.
Reach for it when a run is recorded in one shape and the tool you want to feed expects another, for example scoring a production trace with a chat-based judge, or exporting a chat transcript as a trajectory. It does not fetch traces (see Trace finder) and it does not run agents.
The four formats¶
The same run has four natural records. Each answers a different question, so each keeps different things.
| Format | Class | Unit of data | Keeps | Drops |
|---|---|---|---|---|
| Chat | ChatConversation | One Message (role, content, tool_calls) | What was said, tool calls linked to results by id | Reasoning, timing, model, tokens, cost |
| OpenResponses | ResponsesConversation | One item (message, function_call, function_call_output, reasoning) plus an optional Response per model call whose output holds the output items that call produced | Reasoning items, usage per call, status and error | Timing between items, nesting |
| OTel GenAI trace | OtelTrace | One OtelSpan (an operation with start, end and attributes) | Timing, parent and child tree, per-span usage and cost, subagent nesting | A trajectory-level view; tool results with no call id have no place to live |
| ATIF trajectory | AtifTrajectory | One step (system, user or one agent turn) | Reasoning, tool call and result in the same step, per-step metrics, embedded subagents, audio (v1.8) | Span durations and non-LLM operations |
A chat list is the model's own input, the shape most evaluators and simulators already speak. OpenResponses is the open spec behind the OpenAI Responses API. An OTel trace is what observability tools such as Orq store. ATIF (Agent Trajectory Interchange Format, Harbor's standard) is a turn-by-turn record meant to be replayed, scored or trained on.
Converting: ATIF is the hub¶
Every conversion between two different formats exists as a to_* method on the source object. ATIF is the hub: Responses and OTel convert to and from it directly, and chat reaches it through Responses.
flowchart LR
Chat["ChatConversation"] <--> Responses["ResponsesConversation"]
Responses <--> ATIF(("AtifTrajectory"))
ATIF <--> OTel["OtelTrace"] Chat and Responses also convert directly to each other. Every other pair composes through ATIF, so ChatConversation.to_otel() is to_atif().to_otel(), and it loses whatever both legs lose.
| From To | Chat | Responses | OTel | ATIF |
|---|---|---|---|---|
| Chat | none | to_responses() | to_otel() | to_atif() |
| Responses | to_chat() | none | to_otel() | to_atif() |
| OTel | to_chat() | to_responses() | none | to_atif() |
| ATIF | to_chat() | to_responses() | to_otel() | none |
to_atif() and to_otel() on a non-ATIF source take agent_name and agent_version, which fill AtifTrajectory.agent. OtelTrace.to_atif() falls back to the gen_ai.agent.name and gen_ai.agent.version attributes of the root invoke_agent span when you leave them unset. Both default to 'unknown' otherwise. ChatConversation.to_atif() and ResponsesConversation.to_atif() also take session_id; left unset, it is a hash of the conversation, so the same input always gets the same id.
Add trajectory tags to an evaluation¶
signal_evaluators() turns trajectory tags into evaluatorq scorers. Use it when your output is ATIF, chat, Responses, OTel, or a chat message list and you want deterministic structural tags without an LLM judge. These tags describe execution patterns, not answer quality.
from evaluatorq.signals import signal_evaluators
evaluators = signal_evaluators() # every trajectory tag, bundled thresholds
print([evaluator['name'] for evaluator in evaluators])
Add the returned evaluators to an evaluatorq() run's evaluators list. Unsupported output shapes and signals without enough evidence produce an inconclusive score (value=None, pass_=None) with an explanation, so a missing transcript is visible instead of looking like a clean run.
What each conversion loses¶
Every conversion is lossy in some direction. A dropped tool result, message, subagent or history block logs a loguru warning. OTel spans with operations other than chat, execute_tool and invoke_agent are ignored without a warning. Metadata that has no slot in the target is also dropped without a warning; Field mapping shows where each field lands in every format, and each method's Lost: list names the rest.
- Chat drops reasoning.
ResponsesConversation.to_chat()andAtifTrajectory.to_chat()count the reasoning items they discard and warn once. Keep the Responses or ATIF object if you need the reasoning. - Responses items alone merge consecutive agent steps. Converting to ATIF closes an agent step at a tool result, and at the first output item of each
Responsewhenresponsesare set. Without responses, two assistant turns with no tool result between them become one step.ResponsesConversationchecks that eachResponse.outputmatches the supported model output items (assistant messages, reasoning, function calls, custom tool calls and MCP calls) ofitems, in order, and raisesValueErrorwhen they disagree; aResponsebuilt with an emptyoutputgets it filled fromitems, one response per run of consecutive output items. Compaction is a transcript marker, so aResponse.outputcontaining it raises a clear error. Custom tool and MCP calls and custom tool results stay in ATIF stepextrafor the return trip because ATIF has no matching typed call. - ATIF to Responses emits one
Responseper agent step, with that step's output items as itsoutput, so agent steps stay apart on the way back. Per-call metadata (model, usage, timing) comes from that step; a step without any gets a placeholder that reads back as absent. Responses usage has no unset token count, so an unset one is written as 0 and listed inResponse.metadata['atif_unset_usage'], which reads back as unset. Step cost, token ids and logprobs have no Responses slot and are dropped with a warning. - A call's outcome survives every leg. A step whose model call failed (an
error_type, an OTel span status oferror, or anerrorfinish reason) becomes a Response withstatus: failedand anerror. Alengthorcontent_filterfinish reason becomesstatus: incompletewith the matchingincomplete_details.reason. Responses error codes and incomplete reasons are closed lists, so an error type such astimeoutis written as codeserver_error, and a finish reason such asstophas no field at all. Whatever the Response fields cannot hold exactly goes toResponse.metadata['atif_error_type']and['atif_finish_reasons'], which win when the Response is read back. Reading a Response sets the step'serror_typeandfinish_reasons, so a failed Response also becomes anerrorchat span in OTel. - ATIF to OTel drops results without a call id and unembedded subagents. An observation result with no
source_call_idhas no tool message orexecute_toolspan to sit in, and a subagent reference whose trajectory is not embedded has nothing to render. Each logs a warning. User and system steps are all kept: a system step that opens the trajectory becomes the chat spans'system_instructions, a later one is asysteminput message, and steps after the last agent turn go into a finalchatspan that has input and no output. A trajectory with no agent step at all, such as a system prompt and one user message, converts the same way. - Multi-part text joins with newlines. When one message has several text parts and the target holds a single string, every route joins them with
\nand skips empty parts, soto_chat()andto_atif().to_chat()give the same text. A tool result with several text parts follows the same rule, soaandbbecomea\nbin Responses and OTel. Image and file parts in a Responses tool result stay typed in chat; a malformed optional media field is dropped while a valid source remains. Structured tool result lists become JSON text. Image URLs stay typed in ATIF, while files become warned text markers because ATIF has no file part. A Responses text part whosetextis not a string is dropped with a warning rather than written as a Python repr. - OTel to ATIF reads cumulative and per-turn input. New user, system and assistant messages become steps when a chat span repeats the full history or holds only the current turn. Edited or truncated history logs a warning naming the span; matching turns are aligned by occurrence and include tool arguments and non-text parts. Unmatched messages on either side of a replayed block remain new steps. A changed, nonempty
system_instructionsvalue becomes a new system step. A compaction (an OTelcompactionpart or a Responsescompactionitem) becomes a system step withextra.context_management = {"type": "compaction", "boundary": "replace"}and its original parts underextra["evaluatorq.compaction"]; a Responses compaction item is restored on the return trip to Responses. The first chat span's input is the conversation so far, so its earlier assistant turns become agent steps. Extra output choices are dropped with a warning. Spans that are notchat,execute_toolorinvoke_agentare ignored. - OTel tool results link by call id, never by position. A result in a chat input goes to the call with the same id: one issued earlier in the same input, else one issued by an earlier chat span. A result whose id matches no call, or that has no id, is kept on the nearest agent step with no
source_call_idand its id inextra.orphan_call_id, and logs a warning. Such a result is dropped again by ATIF to OTel, since no tool message holds it there. - Tool results with no
tool_call_idare unlinked. Chat to Responses warns and cannot attach them.
The per-method Lost: list in the API reference is the exact inventory for each edge.
Free-form slots¶
When a source field has no typed home in the target, the converters keep it in the target's free-form slot instead of dropping it. ATIF keeps it in extra on the trajectory, step, tool call or metrics. OTel keeps it in a span's attributes. Reasoning token counts, response status and errors, and fc_ item ids land there when converting Responses to ATIF. ATIF to OTel writes step, trajectory, tool-call and observation-result extra, and final_metrics, as JSON-string span attributes under evaluatorq.atif. (evaluatorq.atif.step.extra on the chat span, evaluatorq.atif.trajectory.extra and evaluatorq.atif.final_metrics on the invoke_agent span, evaluatorq.atif.tool_call.extra and evaluatorq.atif.result.extra on the execute_tool span), and OTel to ATIF reads them back. Additional observation results for one tool call use evaluatorq.atif.additional_results on its execute_tool span. Tool-result start and end timestamps and status map to the corresponding execute_tool span fields, and tool definitions map to gen_ai.tool.definitions on chat spans. Where the trace itself yields a key, such as a step's invocation timing, the trace's value wins and a differing carried value logs a warning. Nothing else reads the evaluatorq.atif.* keys, so treat them as round-trip storage, not as an interchange contract.
Field mapping¶
Each row is one piece of an agent run and where every format keeps it. none means the format has no slot: converting into it drops the value. A value that a converter renames or moves into a free-form slot shows that slot.
| What | Chat | Responses | ATIF | OTel |
|---|---|---|---|---|
| User or system message | Message(role='user') or 'system' | message item with the same role | step with source='user' or 'system' | chat span input message; a system step that opens the run becomes gen_ai.system_instructions on every chat span |
| Developer message | role='developer'; to_chat() renames it to system (warned) | message item with role='developer' | system step with extra.original_role='developer' | input message with role='developer' |
| Assistant text | role='assistant', content | message item with output_text parts | agent step message | chat span output message, text parts |
| Reasoning | none (dropped, warned) | reasoning item | step.reasoning_content | reasoning part of the output message |
| Tool call | tool_calls[]: id, function.name, function.arguments (JSON text) | function_call: call_id, name, arguments | tool_calls[]: tool_call_id, function_name, arguments (a dict) | output tool_call part: id, name, arguments; execute_tool span gen_ai.tool.call.id |
| Arguments that are not a JSON object | kept as the text | kept as the text | arguments={}, the text in the tool call's extra['evaluatorq.raw_arguments'] | gen_ai.tool.call.arguments holds the text |
| Tool result | role='tool': tool_call_id, content, name | function_call_output: call_id, output | observation.results[]: source_call_id, content | execute_tool span gen_ai.tool.call.result, and a tool_call_response part in the next chat input |
| Tool error | none | none | result extra.error_type and extra.status | execute_tool span status='error' and error.type |
| Image | input_image part | input_image part | image part, from image_url only (detail dropped) | uri part; read back into step extra.non_text_parts |
| File | input_file part | input_file part | text marker [file: <name>] | text |
| Audio | none | text marker | audio part (ATIF v1.8) | uri part |
| Model | none | Response.model | step.model_name | gen_ai.response.model; a different gen_ai.request.model goes to extra.requested_model |
| Token usage | none | Response.usage: input_tokens, output_tokens, cached_tokens, reasoning_tokens | metrics: prompt_tokens, completion_tokens, cached_tokens, extra.reasoning_tokens | gen_ai.usage.input_tokens, output_tokens, cache_read.input_tokens, reasoning.output_tokens |
| Cost | none | none (dropped, warned) | metrics.cost_usd | gen_ai.usage.total_cost |
| Time | none | Response.created_at | step.timestamp, and start and end in extra.invocation | span start_time and end_time |
| Finish reason | none | status: incomplete with incomplete_details.reason (max_output_tokens for length, content_filter); any other value in metadata['atif_finish_reasons'] | extra.finish_reasons | gen_ai.response.finish_reasons; the output message's finish_reason holds the first |
| Failed model call | none | status: failed and error.code; an error type that is not a Responses code is written as server_error, with the original in metadata['atif_error_type'] | extra.error_type | chat span status: error and error.type |
| Response id | none | Response.id | extra.response_id | inside the evaluatorq.atif.step.extra attribute |
| Subagent | none | none (dropped, warned) | subagent_trajectories, referenced by a result's subagent_trajectory_ref | nested invoke_agent span under the calling execute_tool span |
| Compaction | none (skipped) | compaction item | system step with extra.context_management | read from a compaction part; not written |
| Agent name and version | none | none; pass agent_name and agent_version to to_atif() | agent.name, agent.version | invoke_agent span gen_ai.agent.name, gen_ai.agent.version |
| Tool definitions | none | none (Response.tools is written empty) | agent.tool_definitions | chat span gen_ai.tool.definitions |
| Session id | none | none; a hash of the items on to_atif() | session_id | invoke_agent span gen_ai.conversation.id, else the trace id |
Field names in the Responses column follow the OpenResponses spec, and those in the OTel column follow the GenAI semantic conventions. ATIF to Responses drops a step's invocation start and end times without a warning, so a run that went OTel to ATIF to Responses keeps them only in the ATIF object.
ATIF versions¶
Read an ATIF document with AtifTrajectory.from_json(text), which takes a string, bytes or an already-parsed dict. It accepts ATIF-v1.7 and ATIF-v1.8; any other schema_version, or none at the document root, raises ValueError, and so does an ATIF-v1.7 document that contains an audio content part, embedded subagents included. Use from_json rather than model_validate_json, which checks neither. A trajectory built in code, and a subagent embedded in a document, may leave schema_version out and get ATIF-v1.7. The two versions differ only by the audio content part.
Writing keeps the trajectory's own schema_version, which is ATIF-v1.7 for one built by a converter. to_json() upgrades to ATIF-v1.8 when any content part in the document, embedded subagents included, is audio. Pass version='1.7' or version='1.8' to force one. Forcing 1.7 on a trajectory that holds audio raises ValueError instead of writing a document that older readers would reject. The chosen version is stamped on the whole document, embedded subagents included.
Deterministic ids¶
The converters never call uuid4. Trace ids, span ids and call ids they have to invent are SHA-256 hashes of a seed: the source's session id or trajectory id when it has one, otherwise a hash of the whole document. Converting the same input twice gives the same output, so a diff between two conversions shows a real change and not fresh random ids.
Real Orq exports¶
OtelTrace.from_orq(spans) takes the raw list of span dicts as Orq returns it (the v3spans list) and parses it into typed spans. It handles what real exports look like rather than what the semantic conventions describe: chat-shaped messages keyed by index, nested singular or plural message wrappers, the chat-completion operation name (read as chat), direct gen_ai.input and gen_ai.output messages on a root span, no execute_tool spans, and usage that lives only in the span summary. Orq's agent role is read as assistant. When a router trace root repeats the same chat input and output as one of its direct children, conversion reads that child once; other children remain separate calls.
Attribute values stay whole: supported message content is parsed into typed parts rather than flattened, so nothing in valid messages is lost on parse. Malformed messages or parts are skipped with a warning.
Examples¶
Each snippet is self-contained. They use plain Python objects, so no API key is needed.
Chat and Responses¶
from evaluatorq.contracts import FunctionCall, Message, StrategyToolCall
from evaluatorq.formats import ChatConversation
chat = ChatConversation(
messages=[
Message(role='developer', content='Answer in one line.'),
Message(role='user', content='Weather in Paris?'),
Message(
role='assistant',
tool_calls=[
StrategyToolCall(id='call_1', function=FunctionCall(name='get_weather', arguments='{"city": "Paris"}'))
],
),
Message(role='tool', tool_call_id='call_1', content='{"temp": 18}'),
Message(role='assistant', content='18C and sunny.'),
]
)
responses = chat.to_responses()
print(responses.items[2]) # {'type': 'function_call', 'call_id': 'call_1', 'name': 'get_weather', 'arguments': '{"city": "Paris"}'}
back = responses.to_chat()
print([message.role for message in back.messages]) # ['system', 'user', 'assistant', 'tool', 'assistant']
print(back.messages[3].name) # get_weather
The return trip renames developer to system and logs a warning. The tool message gets its name back from the call with the same id.
Chat to ATIF¶
import json
from evaluatorq.contracts import FunctionCall, Message, StrategyToolCall
from evaluatorq.formats import ChatConversation
chat = ChatConversation(
messages=[
Message(role='user', content='Weather in Paris?'),
Message(
role='assistant',
tool_calls=[
StrategyToolCall(id='call_1', function=FunctionCall(name='get_weather', arguments='{"city": "Paris"}'))
],
),
Message(role='tool', tool_call_id='call_1', name='get_weather', content='{"temp": 18}'),
Message(role='assistant', content='18C and sunny.'),
]
)
trajectory = chat.to_atif(agent_name='weather-bot', agent_version='1.0')
print([(step.step_id, step.source) for step in trajectory.steps]) # [(1, 'user'), (2, 'agent'), (3, 'agent')]
print(json.loads(trajectory.to_json())['schema_version']) # ATIF-v1.7
ATIF to chat¶
from evaluatorq.formats import AtifTrajectory
trajectory = AtifTrajectory.from_json({
'schema_version': 'ATIF-v1.7',
'agent': {'name': 'demo', 'version': '1'},
'steps': [
{'step_id': 1, 'source': 'user', 'message': 'Hi'},
{'step_id': 2, 'source': 'agent', 'message': 'Hello.'},
],
})
chat = trajectory.to_chat()
print([message.role for message in chat.messages]) # ['user', 'assistant']
For a Harbor run or another saved ATIF document, pass the file's text to AtifTrajectory.from_json instead of the dict.
Responses to OTel¶
from evaluatorq.formats import ResponsesConversation
responses = ResponsesConversation(
items=[
{'type': 'message', 'role': 'user', 'content': [{'type': 'input_text', 'text': 'Weather in Paris?'}]},
{'type': 'function_call', 'call_id': 'call_1', 'name': 'get_weather', 'arguments': '{"city": "Paris"}'},
{'type': 'function_call_output', 'call_id': 'call_1', 'output': '{"temp": 18}'},
{'type': 'message', 'role': 'assistant', 'content': [{'type': 'output_text', 'text': '18C and sunny.'}]},
]
)
trace = responses.to_otel(agent_name='weather-bot')
print([span.operation for span in trace.spans]) # ['invoke_agent', 'chat', 'execute_tool', 'chat']
Orq spans to ATIF¶
from evaluatorq.formats import OtelTrace
raw_spans = [
{
'span_id': 's1',
'parent_span_id': None,
'name': 'chat-completion',
'started_at': '2026-04-20T10:00:00Z',
'ended_at': '2026-04-20T10:00:02Z',
'attributes': {
'gen_ai.operation.name': 'chat-completion',
'gen_ai.response.model': 'openai/gpt-5.6-luna',
'gen_ai.input.messages': [{'role': 'user', 'parts': [{'type': 'text', 'content': 'Hi'}]}],
'gen_ai.output.messages': [{'role': 'assistant', 'parts': [{'type': 'text', 'content': 'Hello.'}]}],
},
}
]
trajectory = OtelTrace.from_orq(raw_spans).to_atif()
print([(step.source, step.message) for step in trajectory.steps]) # [('user', 'Hi'), ('agent', 'Hello.')]
Orq router roots can also hold direct gen_ai.input and gen_ai.output messages, with a child span carrying the same call. A root span's type: trace identifies this wrapper shape:
from evaluatorq.formats import OtelTrace
input_messages = [{'role': 'user', 'parts': [{'type': 'text', 'content': 'Hi'}]}]
output_messages = [{'role': 'agent', 'parts': [{'type': 'text', 'content': 'Hello.'}]}]
router_spans = [
{
'span_id': 'root',
'type': 'trace',
'attributes': {
'gen_ai.operation.name': 'chat',
'gen_ai.input': {'messages': input_messages},
'gen_ai.output': {'messages': output_messages},
},
},
{
'span_id': 'call',
'parent_span_id': 'root',
'type': 'span.responses',
'attributes': {
'gen_ai.operation.name': 'chat',
'gen_ai.input.messages': input_messages,
'gen_ai.output.messages': output_messages,
},
},
]
trajectory = OtelTrace.from_orq(router_spans).to_atif()
print([(step.source, step.message) for step in trajectory.steps]) # [('user', 'Hi'), ('agent', 'Hello.')]
A trace with no chat span that yields a step raises ValueError from to_atif(), so check the span list is not empty or tool-only before converting.
An OtelTrace holds one trace. Spans that carry more than one distinct trace_id raise ValueError when the trace is built, by from_orq or directly, so group spans by trace id first. Spans with no trace id are accepted, and several root spans under one trace id are converted as one run in start-time order.