Skip to content
AI News HubLIVE
Source content · Analysis pending6 min read

AI Agent Observability: Logging, Tracing, and Debugging Explained

Summary

Chain Visualization: Reading the Trace Waterfall The spans from the last section don't mean much as a raw list.

SourceMachine Learning MasteryAuthor: Shittu Olumide
AI Agent Observability: Logging, Tracing, and Debugging Explained
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

AI Agent Observability: Logging, Tracing, and Debugging Explained - MachineLearningMastery.com AI Agent Observability: Logging, Tracing, and Debugging Explained - MachineLearningMastery.com In this article, you will learn what AI agent observability means, why traditional monitoring tools fall short for agentic systems, and how to implement structured logging, distributed tracing, and practical debugging workflows for AI agents. Topics we will cover include: Why AI agents fail in ways that look like success, and why that makes standard monitoring tools insufficient. How to implement structured logging and OpenTelemetry-based tracing for agent runs, tool calls, and model inference steps. How to read trace waterfalls, track token costs with metrics, and use that data to debug real agent failures. An agent handling customer support tickets closes one out with a clean, professional, entirely wrong answer. It called the refund-lookup tool once, then called it again with slightly different arguments a few seconds later, then answered confidently based on the second result instead of the first. Nothing crashed. No error fired. The uptime dashboard shows green the entire time. The only reason anyone finds out is a customer replying two days later, confused, and by then nobody can reconstruct what actually happened inside that run. That failure is the whole reason this article exists. A traditional service either returns an error 200 or throws something you can grep for. An agent can do neither and still be completely wrong, and the tooling built for the first kind of system is close to blind to the second. This article walks through what actually needs to change — logging, tracing, and debugging — one at a time, with real code. What AI Agent Observability Actually Means AI agent observability is the practice of capturing every model call, tool execution, and reasoning step an agent makes as structured data, so that when something goes wrong, you can reconstruct exactly what happened and why, rather than guessing or re-running the same prompt and hoping the problem repeats itself. It borrows from the three pillars observability engineers already know — logs, metrics, and traces — but the reason it needs its own name and its own discipline comes down to how agents actually fail. Aryan Kargwal, a researcher in the field, put it plainly in coverage from Digital Applied’s 2026 observability guide: agentic systems fail in ways that look like success — incorrect but well-formed outputs, unnecessary tool calls, or actions that are syntactically valid but semantically wrong. None of that trips an error handler. A health check reporting “up” tells you almost nothing useful about whether the agent actually did the right thing on any given run. Why Agents Break the Traditional Monitoring Model It’s worth being specific about the mechanics here, because “agents are unpredictable” undersells exactly what changes. The same input doesn’t reliably produce the same behavior anymore. Temperature settings, retrieval results, and which tools happen to be available can all shift the path an agent takes, so the same prompt can trigger a genuinely different sequence of tool calls on two consecutive runs. A single “it worked when I tested it” trace tells you almost nothing about what the distribution of real runs actually looks like. Cost and latency stop correlating with request count and start correlating with tokens instead. A single “slow” request might be consuming ten times the normal token budget, and a monitoring setup built around requests-per-second is structurally blind to that. Multi-step chains compound the problem: one user request might trigger several model calls, a handful of tool calls, and a couple of retrieval lookups, and each one is an independent point of failure that a single aggregate error metric can’t distinguish between. And prompts themselves routinely carry real personal or confidential information, which means naive logging that dumps full prompt text into a backend creates a genuine compliance problem before it has created any debugging value at all. A side-by-side comparison of the two worlds makes the shift concrete: Signal Traditional app LLM / AI agent Latency driver CPU, I/O, network Token count, model size, context window Cost unit Requests per second Tokens consumed Failure mode Exception, timeout Hallucination, context overflow, tool error Debug artifact Stack trace Prompt, completion, and the reasoning chain between them Logging Start with the most familiar pillar, because it’s still the foundation everything else builds on, just applied differently. For an agent, the events worth logging are specific: which tool got called and with what arguments, what came back, how many tokens a given step consumed, how long each hop took, and any error along the way — and all of it structured rather than written as free-text sentences a human has to parse later. The detail that actually makes agent logging useful is tying every log line back to the specific run it came from. A log statement that just says “tool call failed” is nearly worthless at 2 am when three different users triggered three different runs in the same minute. Attaching the current trace ID to every log line — something OpenTelemetry does automatically once tracing is set up — is what turns a pile of scattered log statements into something you can filter down to the exact run that broke. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 import logging from opentelemetry import trace # Standard Python logging, nothing exotic here logger = logging.getLogger("agent") logging.basicConfig(level=logging.INFO) tracer = trace.get_tracer("agent-service") def call_tool(tool_name: str, arguments: dict): # get_current_span() pulls whatever span is active right now, # so the log line below can be tied back to the exact trace # and step it happened inside span = trace.get_current_span() trace_id = format(span.get_span_context().trace_id, "032x") logger.info( "tool_call_started", extra={ "trace_id": trace_id, "tool_name": tool_name, "arguments": arguments, }, ) try: result = execute_tool(tool_name, arguments) logger.info( "tool_call_succeeded", extra={"trace_id": trace_id, "tool_name": tool_name, "result_length": len(str(result))}, ) return result except Exception as e: logger.error( "tool_call_failed", extra={"trace_id": trace_id, "tool_name": tool_name, "error": str(e)}, ) raise A few things worth noticing in that snippet. trace.get_current_span() doesn’t require you to manually pass a trace ID down through every function call — it reads whatever span is active in the current execution context, which is exactly what makes this pattern practical to sprinkle throughout a real codebase without threading an ID parameter through every layer. Logging the arguments and the result length, rather than the full result content, is a deliberate choice, not an oversight; full tool outputs can be large and can carry sensitive data, and a length or a truncated preview is usually enough to spot a problem without turning every log line into a privacy liability. And logging both a start and an end event for the same tool call, rather than just the outcome, is what lets you later measure exactly how long that specific call took — which becomes the raw material tracing formalizes properly in the next section. Tracing Logging tells you what happened at individual points in time. Tracing is what stitches those points into a shape — a full record of one agent run from the first request to the final answer, with every step nested inside the step that triggered it. That nested shape is the actual answer to “why did the agent do that,” because it shows you not just that a tool was called, but which reasoning step decided to call it and what happened immediately before and after. The vocabulary here comes from the OpenTelemetry GenAI semantic conventions, which define a standard set of gen_ai.* span types and attributes specifically for this. Rather than every team inventing their own span names, the spec defines a handful of operation types worth knowing: create_agent for when an agent is first defined, invoke_agent for a single agent run, invoke_workflow for orchestration across multiple agents handing off to each other, execute_tool for an individual tool call, and chat for the actual model inference call itself. Each one carries a standard set of attributes — gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, and gen_ai.response.finish_reasons among them — so a trace produced by one team’s agent looks structurally the same as one produced by a completely different framework. Here’s what manually instrumenting a small tool-calling agent actually looks like: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 from opentelemetry import trace from opentelemetry.trace import Status, StatusCode tracer = trace.get_tracer("agent-service") def run_agent(task: str) -> str: # The root span for this entire run; every step below nests # inside it, which is what produces the parent-child tree with tracer.start_as_current_span("invoke_agent") as agent_span: agent_span.set_attributes({ "gen_ai.system": "openai", "agent.name": "support-agent", "gen_ai.request.model": "gpt-4o", }) messages = [ {"role": "system", "content": "You are a support assistant."}, {"role": "user", "content": task}, ] while True: # The model call itself gets its own child span with tracer.start_as_current_span("chat") as chat_span: response = model_client.chat.completions.create( model="gpt-4o", messages=messages, tools=AVAILABLE_TOOLS ) choice = response.choices[0] chat_span.set_attributes({ "gen_ai.response.model": response.model, "gen_ai.usage.input_tokens": response.usage.prompt_tokens, "gen_ai.usage.output_tokens": response.usage.completion_tokens, }) if choice.finish_reason != "tool_calls": agent_span.set_status(Status(StatusCode.OK)) return choice.message.content # Each tool call gets its own child span, nested under # the agent run, not under the chat span, since a tool # call is a sibling step, not a sub-step of inference for tool_call in choice.message.tool_calls: with tracer.start_as_current_span("execute_tool") as tool_span: tool_span.set_attributes({ "gen_ai.tool.name": tool_call.function.name, "gen_ai.tool.call.id": tool_call.id, }) try: result = call_tool(tool_call.function.name, tool_call.function.arguments) except Exception as e: tool_span.record_exception(e) tool_span.set_status(Status(StatusCode.ERROR, str(e))) raise messages.append({ "role": "tool", "content": str(result), "tool_call_id": tool_call.id, }) The nesting is doing the real work here. Every chat span and every execute_tool span opens inside the with tracer.start_as_current_span(…) block belonging to the run above it, which is exactly what OpenTelemetry uses to build the parent-child relationship automatically — you never manually wire “this span belongs under that one,” it’s implicit in how the with blocks are structured in your code. record_exception plus set_status(StatusCode.ERROR, …) on the tool span is what makes a failed tool call show up clearly in a trace viewer rather than silently vanishing into the returned string, which matters directly for the debugging section later in this article. And separating token usage attributes onto the chat span specifically, rather than the top-level invoke_agent span, is what lets a trace viewer later show you token cost broken down per model call within a single run, not just a single combined total for the whole thin [truncated for AI cost control]

Key points and analysis

Article intelligence

EngineersIntermediate

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Chain Visualization: Reading the Trace Waterfall The spans from the last section don't mean much as a raw list.

Highlights and analysis are generated automatically and may contain errors. Check the original source.