October 2, 2026
What to Instrument in an AI Agent: A Practical Guide to Traces, Spans, and Agent-Level Observability

Contents
Key Takeaways
- AI agent observability goes beyond APM: it records what an agent decided, what it based that decision on, and whether the action was correct.
- The same telemetry serves debugging, cost attribution, evaluation, and audit, so gaps cost twice: once in the investigation and again in the review.
- Instrument six layers: prompts and completions, tool calls, retrieval context, policy decisions, cost and latency, and the identity each call ran under.
- Mask sensitive data before storage, and set retention by data type. For high-risk systems under the EU AI Act, six months is the minimum.
- In multi-agent systems, pass trace context across every handoff so a failure can be followed back to its cause.
- Keep every anomalous trace, register each agent with a stable identity, and link traces to their evaluations and policy decisions.
An AI agent can complete every step of a task without a single error and still get the outcome wrong. Imagine an agent processing a customer refund: the model responds, the tool call succeeds, and the payment API confirms the transaction. Everything looks fine, except the money went to the wrong account. Traditional monitoring wouldn’t flag this, because it checks whether a system is working, not whether it made the right decision. Closing that gap starts with instrumentation, which means adding the code and tooling that record what an agent does at each step.
What you record matters to more than one team. Engineers rely on it to debug failures, while validators and auditors treat it as evidence. This guide covers what to capture at each layer, what to mask, how long to keep it, and what to check before your next agent ships.
How Is AI Agent Observability Different from APM?
Application performance monitoring (APM) is the set of tools teams use to track whether software is up, fast, and error-free. AI agent observability goes further: it shows what an agent decided, what it based that decision on, and whether the resulting action was correct. It also extends LLM observability, which focuses on individual model calls, to the chain of decisions and actions that connects them. That chain is what agentic AI monitoring needs to cover.
The table below compares APM and agent observability.
| APM | AI agent observability | |
| Core question | Is the system working? | Did the agent make the right decision? |
| What it measures | Latency, error rates, throughput, and uptime | Prompts, retrieved context, tool parameters, policy decisions, and outcomes |
| Behavior it assumes | Deterministic: the same input produces the same output | Non-deterministic: the same input can produce different outputs |
| What counts as a failure | An error, timeout, or crash | A wrong or unsafe action, even when every call succeeds |
| What a trace contains | Timing and status of each request | Timing and status, plus content, context, identity, and the policy applied |
| Who relies on it | Engineering and operations teams | Engineering, AI Ops, and risk, compliance, and validation teams |
The most important difference is in how failure shows up. Most agent failures never register as errors: the API call succeeds, even when the parameters inside it were wrong. This is why agent observability keeps APM’s building blocks, such as traces and spans, but changes what goes inside them.
For risk and compliance teams, the takeaway is simple. A clean error log is not evidence that an agent behaved correctly.
The Core Building Blocks: Trace, Span, Tool Call, Session, and Agent
Agent observability rests on five terms. Together, they describe how an agent’s activity is recorded, from one request and the steps inside it, up to the agent itself. Knowing them makes it easier to decide what to capture, and where.
Trace. A trace is the complete record of one request, from the input an agent receives to the final output or action it produces. It is made up of spans that share a single trace ID. When something goes wrong, the trace is where an investigation starts.
Span. In LLM tracing, a span is a single, timed unit of work inside a trace, such as a model call, a retrieval step, or a tool call. Each span records a start and end time, a status, and attributes that describe what happened. Spans also point to a parent span, which is how a trace shows the order of steps and which step triggered which.
Tool call. A tool call is a span that records the agent acting on an external system, such as an API, a database, or a payment service. It captures the tool’s name, the parameters passed, and the result. This is where an agent’s decisions turn into real consequences.
Session. A session groups the traces that belong to one conversation or multi-step task. A single trace shows one request. A session shows how context builds across requests, which matters when an early answer shapes a later action. Session replay lets a reviewer step through those traces in order and see what the agent knew before each action.
Agent. The agent is the stable identity that sessions and traces roll up to. With a consistent name and version, teams can compare an agent’s behavior over time, attribute cost, and pull every record tied to it. Without one, traces remain isolated events.
These terms are not unique to any one vendor. OpenTelemetry, the open-source standard for collecting telemetry, publishes semantic conventions for generative AI. If you use OpenTelemetry for LLM and agent tracing, these conventions define standard span types, such as invoke_agent for an agent run and execute_tool for a tool call. They also define common attribute names for the model used, token counts, and the agent’s identity. The conventions now live in their own repository and are still in development, but following them keeps your telemetry portable across tools.
Why Agent Observability Is an Evidence Problem, Not a Telemetry Problem
Picture the review after an incident. A validator asks what data the agent retrieved, which policy it checked, and whose permissions it used when it acted. If the answer is “we didn’t record that,” the investigation stops there. So does the audit. This is why agent observability is an evidence problem. Telemetry is usually treated as an engineering concern, collected to keep systems running. For AI agents, the same records serve four jobs at once:
- Debugging: finding where and why an agent went wrong
- Cost attribution: tracing spend back to a specific agent, task, or team
- Evaluation : scoring outputs against quality and policy criteria
- AI audit: showing reviewers and regulators what the agent did, and why
When instrumentation is partial, the cost is paid twice. First, the incident can’t be fully investigated, so the root cause stays hidden and the failure can repeat. Second, the evidence isn’t there when a validator or examiner asks for it. Adding logging afterward doesn’t close the gap, because it can’t recover events that were never captured.
Regulation points the same way. Under the EU AI Act, high-risk AI systems must technically allow for the automatic recording of events over their lifetime, and deployers must keep those logs for at least six months. Following the Digital Omnibus, in force since July 2026, these obligations apply from 2 December 2027 for high-risk systems in areas such as hiring and credit scoring. In the US, banking regulators left generative and agentic AI out of their revised model risk guidance, but plan to seek input on how banks use them. Teams that capture complete records now won’t have to rebuild them under pressure later.
What Should You Log for an AI Agent? A Layer-by-Layer Guide
Each layer below follows the same pattern. The first part covers what to capture and what to mask, for the teams building the agent. The second covers why an investigation fails without it and how long to keep it, for the teams that will review it.
Prompts and Completions
Capture: the user input, the model’s output, the model name and version, and settings such as temperature. For the system prompt, store a version ID so you can always tell which instructions were live. Mask personal data, credentials, and account numbers before anything is written to storage, not afterward. OpenTelemetry’s GenAI conventions treat capturing prompt and response content as optional, so recording it is a deliberate decision your team needs to make.
Why it matters: without the prompt and output, reviewers can see that an agent acted, but not what it was told or what it said. Prompt injection attempts stay invisible. Full text is the most sensitive data in the trace, so it usually gets the shortest retention.
Tool Invocations and Parameters
Capture: the tool name, the parameters passed, the result returned, and whether the call was allowed, blocked, or modified. When you redact parameter values, keep the field names, so reviewers can still see the structure of the call.
Why it matters: tool calls are where decisions become actions. Without the parameters, you can see that a payment was made, but not to whom or for how much. Keep tool call records for as long as any record of the action they caused.
Retrieval Context
Capture: the query, the IDs or sources of the documents returned, their ranking, and the version of the index searched. Store references rather than full documents where possible. This keeps storage manageable and avoids copying sensitive content into logs.
Why it matters: many wrong answers start with wrong context. Without a record of what was retrieved, you can’t tell whether the model failed or the retrieval step did. Keep the references as long as the decisions they informed, along with the index version, so the context can be reconstructed later.
Policy Decisions
Capture: which policy or guardrail was checked, its version, what it evaluated, and the outcome: allowed, blocked, modified, or escalated. Record any human approval as well, and link each decision to the span it governed.
Why it matters: this is the record that proves controls worked. Without it, you can say a guardrail exists, but not that it fired. It’s often the first thing a validator asks for, so policy decisions should be kept for the full audit period.
Token Counts, Cost, and Latency per Hop
Capture: input and output tokens for each model call, the resulting cost, and the duration of each span, not just the total for the trace.
Why it matters: per-hop detail shows where time and money actually go, such as a retry loop, a slow tool, or an oversized context. Without it, cost can’t be attributed to a specific agent or team, and spikes can’t be explained. Raw per-span figures can be kept for shorter periods, while aggregated totals follow your financial reporting cycle.
The Identity the Call Ran Under
Capture: the agent’s ID and version, the user or service account it acted for, and the permissions it used. Record the scope of credentials, never the credentials themselves.
Why it matters: accountability depends on knowing who acted and on whose behalf. Without identity, you can’t tell whether an agent exceeded its permissions or whose data it touched. Identity fields should be kept for as long as the records they’re attached to.
What a Complete Tool Call Span Looks Like
Here is a single tool call span from the refund scenario in the introduction, using OpenTelemetry attribute names where they exist and masking the sensitive value:
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 |
span: execute_tool issue_refund trace_id: 4bf92f3577b34da6a3ce929d0e0e4736 parent_span_id: 00f067aa0ba902b7 start_time: 2026-09-14T10:42:07.118Z duration_ms: 412 status: OK gen_ai.operation.name: execute_tool gen_ai.tool.name: issue_refund gen_ai.tool.call.id: call_7d21 gen_ai.tool.call.arguments: {"account_id": "[REDACTED]", "amount": 240.00, "currency": "EUR"} gen_ai.tool.call.result: {"refund_id": "rf_5521", "status": "confirmed"} gen_ai.agent.id: refunds-agent gen_ai.agent.version: 2.3.1 app.acted_for: support-portal-service app.permissions: refunds:write (limit 500 EUR) app.policy.name: refund-limit-check app.policy.version: 4 app.policy.outcome: allowed |
Fields that start with gen_ai. follow the OpenTelemetry conventions. The app. fields are custom attributes your team defines for identity and policy details. Even with the account number masked, a reviewer can see what the agent did, under which permissions, and which policy allowed it.
How Long to Keep Agent Telemetry
Retention works best as a tiered decision. Full prompt and response text is the most sensitive, so it usually has the shortest retention. Metadata, policy decisions, evaluation results, and identity fields are smaller and more valuable as evidence, so they are kept longer. Under the EU AI Act, deployers of high-risk systems must keep the automatically generated logs under their control for at least six months, unless other EU or national law, in particular data protection law, provides otherwise. Sector rules can also require longer periods. Masking sensitive fields at capture keeps records useful as evidence without holding more personal data than needed.
A recent case shows why retained records matter. After a similar disclosure by OpenAI, Anthropic reviewed its cybersecurity evaluation transcripts and found three incidents where Claude models, during testing, reached real systems they were never meant to access. The earliest dated back to April 2026. Those incidents came to light only because the records still existed, and Anthropic noted that real-time monitoring of those logs could have surfaced the problem sooner.
How to Debug Multi-Agent Systems
In a multi-agent system, a failure often surfaces several agents away from where it started. One agent retrieves the wrong information, a second summarizes it faithfully, and a third acts on the summary. The third agent looks like the problem, even though it did exactly what it was told.
Debugging multi-agent systems means walking backward from the symptom to the cause. Four practices make that possible:
- Context propagation: each agent passes the trace context to the next, so every step joins the same trace instead of starting a new one.
- A shared trace ID: all spans from one request carry the same ID, so they can be viewed together.
- Parent-child links: each span records which step triggered it, showing the order of events and the path the request took. When agents run asynchronously, span links can connect related traces that don’t share a parent.
- Handoffs recorded as spans: the moment one agent passes work to another is captured as its own span, including what was passed.
Here’s how this works in practice. An underwriting agent approves a policy it should have declined. Its own spans show a reasonable decision based on the summary it received. Following the parent span back leads to the handoff from a summarization agent, and the handoff span shows the summary left out a prior claim. One hop further, the retrieval span explains why: it pulled an outdated customer record.
Without linked traces, that investigation ends at the underwriting agent. For reviewers, the risk is concrete: the wrong agent gets blamed and fixed, while the real cause stays in production.
How to Avoid the Four Most Common Instrumentation Mistakes
Most instrumentation gaps aren’t caused by missing tools. They come from design decisions that seem reasonable at the start and only show their cost during an investigation.
Sampling That Discards the Traces Worth Keeping
Many teams sample traces to control storage costs. The problem is how the decision gets made. With head-based sampling, the system decides whether to keep a trace when the request starts, before it knows whether anything unusual will happen. Rare failures, which are the ones worth investigating, get dropped at the same rate as routine requests. Tail-based sampling decides after the trace completes, so it can keep what matters.
Fix: keep every trace with an error, a blocked or escalated policy decision, or unusual cost or latency, and sample the rest.
Personal Data in Logs and Context Windows
Sensitive data enters an agent from more places than user input. It arrives through retrieved documents and tool results too. Once it lands in the model’s context window, it can be repeated in outputs, passed to other agents, and written to logs, where it spreads to dashboards and exports.
Fix: redact at the point of capture, before storage, and apply the same rules to retrieval results and tool outputs.
No Stable Agent Identity
When agents are identified only by a service name or deployment, their identity can change with every release. Traces can’t be grouped by agent, compared across versions, or tied to an owner. Unregistered agents also become a form of shadow AI, running without anyone accountable for them.
Fix: register each agent with a unique ID, a version, and a named owner, and attach them to every span.
No Link Between a Trace and Its Evaluation or Policy Decision
Evaluation scores often live in one tool, policy logs in another, and traces in a third. A reviewer then can’t see which trace received which score, or which policy governed which action.
Fix: record the trace ID on every evaluation result and policy decision, so a single lookup returns the full picture.
For each evaluation, capture the evaluator and its version, the score, and the criteria it was scored against. OpenTelemetry’s conventions support this with an evaluation result event that attaches to the span it scores. For agent evaluation in production, score a sample of live traces, not just test sets, plus every trace that was flagged, blocked, or escalated. Then feed human review findings back into the evaluation set, so the next round of scoring reflects what reviewers caught.
A Pre-Ship Instrumentation Checklist for AI Agents
Use this checklist before an agent goes into production. The first group is for the team building the agent. The second is for the team that will review or validate it.
For the Team Building the Agent
- Prompts, outputs, model version, and system prompt version are captured, with personal data masked before storage.
- Every tool call records the tool name, redacted parameters, result, and outcome.
- Retrieval spans record document references and the index version.
- Every policy check records the policy, its version, and the outcome, linked to the span it governed.
- Tokens, cost, and latency are captured for each span.
- Trace context carries across every handoff between agents.
- Sampling keeps every trace with an error, a policy hit, or an outlier.
- The agent is registered with an ID, version, and owner, attached to every span.
For the Team Reviewing or Validating the Agent
- The agent has a named owner and a documented purpose.
- Retention periods are set for each type of data and meet legal minimums.
- Every action can be traced to the identity and permissions it ran under.
- Evaluation results and policy decisions carry trace IDs.
- The agent’s full record can be retrieved in one place.
- A test investigation has been run end to end, from a symptom back to its cause.
How Lumenova AI Supports Agent-Level Observability
Lumenova AI builds on the same idea as this guide: one record that serves engineers and reviewers alike. AI observability records every request as a trace broken into spans, so an investigation can follow an agent step by step. Agent registration gives each agent the stable identity described in the mistakes section, so every trace and session rolls up to a dedicated agent page with dashboards, evaluation stats, and topic clustering.
For policy decisions, the Constraint Trace Explorer records the timestamp, tool name, redacted parameters, matched policy, and outcome of each check, which is the record validators typically ask for first. Human annotations feed review findings back into the evaluation cycle, so what reviewers catch improves the next round of scoring. And because invocation history, evidence, cost, evaluations, and alerts sit in one place, no one has to stitch records together from separate systems during an investigation or an audit.
Build the Record Before You Need It
An agent that completes every call successfully can still take the wrong action. Traditional monitoring won’t show the difference, but a well-instrumented trace will.
The teams that get this right treat telemetry as evidence from the start. They capture each layer deliberately, mask what should never be stored, and keep records long enough to answer the questions that come later. The same record then lets engineers find the cause and lets validators confirm it. Instrument before the next agent ships, not after the first incident.
Book a demo to see how Lumenova AI brings traces, evaluations, and policy decisions together for every agent.
Frequently Asked Questions
AI agent observability is the practice of recording what an agent does at each step: what it was asked, what it retrieved, which tools it called, which policies applied, and what it cost. Unlike APM, it shows whether the agent made the right decision, not only whether each call succeeded.
Yes, but only going forward. Instrumentation added later can’t recover events that were never recorded, which is why it’s best done before launch.
A span is a single, timed step inside a trace, such as a model call, a retrieval, or a tool call. Each span records when it started and ended, its status, and attributes such as the model used and token counts. Parent-child links between spans show which step triggered which.
Only when the agent is part of a high-risk AI system, such as one used in hiring or credit decisions. Those systems must record events automatically, and deployers must keep the logs for at least six months. For these systems, the obligations apply from 2 December 2027.
Yes, if sensitive data is masked before storage and access is restricted. Without prompts, many failures can’t be investigated, so the goal is to log them safely rather than skip them.