NVIDIA NeMo Relay now provides a way to inspect the detailed steps an AI agent takes to complete a task, revealing inefficiencies hidden behind successful outcomes. Even a correct final answer can mask problems; for example. Developers can now use NeMo Relay, integrated natively with the Hermes Agent runtime, to examine event streams and trajectories, including model and tool calls, errors and token use, to understand how an agent achieves a result. The technology is demonstrated in a Hermes ToolPerf case study, evaluating harness changes across multiple runs using OpenTelemetry traces explored in Arize Phoenix.
NVIDIA NeMo Relay Observability: ATOF, ATIF, and OpenTelemetry
The potential for subtle inefficiencies in AI agents is now under scrutiny with the introduction of detailed execution tracing, as demonstrated by the Agent Trajectory Interchange Format (ATIF), a step-by-step JSON record of agent interactions and observations. This format assembles data from lifecycle events, allowing developers to move beyond simply verifying a final answer and instead analyze the process by which an agent arrives at that conclusion.
A specific example of this hidden inefficiency is when an agent repeatedly fetches the same content, even if the ultimate result is correct. NVIDIA’s NeMo Relay facilitates this level of inspection by providing two key representations of agent execution alongside ATIF: the Agent Trajectory Observability Format (ATOF), a JSONL log of events with timestamps and IDs for reconstructing the agent run, and the ability to integrate with tools like Arize Phoenix.
Developers can use ATOF to audit individual events, pinpoint timing issues and understand the relationships between parent and child processes within the agent’s workflow. The combination of these formats allows for a granular understanding of how an agent functions, moving beyond a simple pass/fail assessment of its output, Hermes says. The practical application of this technology is demonstrated through a file-and-web research task, where generated OpenTelemetry traces are explored within Arize Phoenix.
This allows for visual analysis of agent behavior and facilitates the evaluation of changes made to an agent’s harness. The entire process is designed to be self-contained, using a local environment created with Docker and specific versions of Python 3. 11, Hermes 0. 21. 1, and NeMo Relay 0. 8. 3, ensuring consistency and reproducibility. The setup script avoids modifying existing Python or Hermes installations, streamlining the testing process.
Running the task generates both ATIF and ATOF trace files, which are then checked by the runner to verify the response. A successful run produces a verification result alongside summaries of the trace data. “The Hermes Agent and NeMo Relay processes run from this repository’s local environment,” according to the project documentation, emphasizing the ease of deployment and analysis. The system’s flexibility extends to compatibility with other OpenTelemetry-compatible backends, such as LangSmith, allowing developers to choose the analysis platform that best suits their needs.
The example task reuses the NVIDIA Nemotron model and requires an NVIDIA API key for access. Hermes uses its built-in keyless web search functionality, and the Phoenix service is launched within a pinned local container. Running the conference search example demonstrates the complete workflow, from task execution to trace generation and analysis.
The companion repository provides all necessary artifacts, including runnable tasks, verifiers, NeMo Relay configurations and Phoenix setup instructions. Detailed documentation is available for NeMo Relay installation, concepts and configuration, as well as for installing and using the Hermes Agent. This comprehensive support aims to empower developers to effectively use these tools for building and debugging AI agents. Further resources include documentation on ATIF and ATOF formats, NeMo Relay observability configuration, and supported integrations with other agent frameworks and applications.
Hermes Agent and NeMo Relay: Integrated Agent Execution Tracking
Inefficiencies in AI agent task completion can be masked by a correct final answer, as a truncated file read can repeatedly fetch the same content, increasing latency and token consumption. While a successful outcome confirms task completion, it does not explain how the agent arrived at that result, or whether the process was optimal. To address this, developers are now using detailed execution traces to dissect agent behavior and pinpoint areas for improvement.
NeMo Relay records lifecycle events, preserving timing and parent-child relationships, and outputs these traces in multiple formats for different analytical needs. Agent Trajectory Observability Format, or ATOF, is a JSONL log of scope starts, scope ends, and point-in-time marks, useful for debugging individual events and examining timing.
Evaluating Agent Harnesses with Trace Evidence and Task Verification
The ability to pinpoint inefficiencies within an AI agent’s reasoning process is now enhanced through detailed tracing, as a truncated file read can lead an agent to repeatedly request the same information, masking performance issues even when the final answer is correct. Developers can now use NeMo Relay to move beyond simply assessing whether an agent succeeds or fails, and instead analyze how it arrives at a solution by inspecting model and tool calls, errors, retries, duration and token usage.
This granular level of inspection is facilitated by the native integration of NeMo Relay with the Hermes Agent runtime, allowing for examination of the complete event stream associated with a task. The process of evaluating agent behavior benefits from the use of Arize Phoenix, a platform that visualizes OpenTelemetry traces generated during task execution, specifically within the file-and-web research task run with NeMo Relay.
This visualization allows for a comparative analysis of agent behavior across multiple runs, as demonstrated in the Hermes ToolPerf case study, and provides a means to assess the impact of modifications to the agent’s underlying harness, according to the company. Successful task completion, such as a simple terminal tool execution printing “VALUE=42”, is an exact success check, verifying that Hermes successfully accessed the model, invoked the terminal tool in a sandboxed environment, and generated both NeMo Relay trace files, ATOF and ATIF.
The system preserves the ordered sequence of lifecycle events through the Agent Trajectory Observability Format, or ATOF, while the Agent Trajectory Interchange Format, or ATIF, presents the same information as a readable trajectory. This pairing of trace formats with a deterministic verifier enables developers to compare harness changes without misinterpreting a reduction in calls as an improvement in performance.
The traces themselves are archived in a results directory, allowing for independent verification and regeneration of performance tables from the raw data records. “NeMo Relay gives you a consistent way to capture evidence for evaluating changes to optimize your agent harnesses,” according to the documentation, highlighting the technology’s focus on quantifiable improvements. The methodology extends to comparing different models on the same task, maintaining consistent queries, tools, execution limits and verification criteria to isolate the impact of the model itself on the execution path.
While live web searches introduce variability, controlled comparisons require fixed search responses and repeated runs to ensure reliable results. The ability to examine the complete agent run within Phoenix, from initial file access to final response, provides a detailed view of the process, including conference queries, file read requests, timing and token usage.
This level of detail allows developers to identify bottlenecks and inefficiencies that might otherwise remain hidden, even when the agent produces a correct final answer. The framework’s utility lies in its ability to capture and analyze the complete execution history, enabling a more detailed understanding of agent behavior and facilitating targeted optimization of the agent harness, the company says.
By pairing these traces with a deterministic verifier, developers can confidently assess the impact of changes without being misled by superficial improvements. The system’s design prioritizes consistency and reproducibility, ensuring that evaluations are based on verifiable evidence rather than subjective impressions.
Hermes Agent Setup: Isolated Runtime with NeMo Relay 0.8.3
Hermes Agent’s initial task run focuses on a deliberately small example, allowing developers to confirm complete setup before incorporating web searches and integration with Arize Phoenix. The system uses a terminal tool executing a Python script within a Docker container to print “VALUE=42”. The container operates in isolation, lacking access to the network, the repository checkout, or the developer’s NVIDIA API key.
Hermes cannot revert to executing terminal commands on the host system. To begin, developers clone the tutorial repository from NVIDIA’s GitHub account and navigate to the examples/tools/hermes-relay-tracing directory; a script then establishes the isolated Hermes Agent and NeMo Relay runtime.
Following this, the keys. env. example file is copied and the user’s NVIDIA API key is added to the keys. env file before proceeding, ensuring proper authentication with the NVIDIA Nemotron 3. 5 Lightning model. Verification that Docker is running is a prerequisite, followed by building a Docker image specifically for the terminal-tool task using a provided script. Finally, executing the run_tutorial.
NeMo Relay functions by providing a standardized method for observing and controlling model and tool execution within an agent; the Hermes Agent harness natively integrates NeMo Relay, representing sessions, turns, model calls, and tool calls within NeMo Relay’s hierarchical structure. This detailed record allows for precise examination of how an agent arrives at a solution, not just whether it succeeds.
Developers can consult NeMo Relay documentation for installation, concepts and configuration, alongside Hermes Agent documentation for installation and usage guidance, the company states. “NeMo Relay gives agent developers a common way to observe and control model and tool execution,” according to the documentation.
Tool-Use and Research Tasks: Verifying Success with NeMo Relay Traces
The Nous Research Hermes ToolPerf benchmark used NeMo Relay ATOF traces to establish ground-truth turn accounting after analyzing production sessions and auditing tool schemas. This benchmark, built from nine identified failure patterns, was then used to evaluate a series of fixes to the Hermes tool layer, demonstrating a method for rigorous assessment beyond simple task completion rates, by the company’s account.
Analysis of the hidden-file search task revealed a persistent gap between the baseline and fixed revisions, remaining at 0 to 33% success on both models, a finding that task success alone would not have revealed. The NeMo Relay traces recorded every model call, tool call, error, retry, result payload, and timing for all 108 runs, allowing for a granular examination of agent behavior. This detailed logging uncovered that Qwen’s additional turns were not indicative of failure, but rather recovery attempts, and traced a regression in one task back to a specific probe output.
The ability to dissect agent behavior at this level of detail is important, as a correct final answer can mask underlying inefficiencies. The combination of these traces with a deterministic verifier allows developers to differentiate between genuine improvements and simply fewer calls, ensuring that optimizations translate to better results, not just reduced resource consumption.
This granular insight extends beyond overall success rates, enabling the identification of specific bottlenecks and areas for improvement within the agent’s decision-making process. The Hermes ToolPerf case study demonstrates how this approach can be applied to evaluate changes to tool layers, providing a robust framework for continuous optimization and refinement of AI agents.




See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
