Build Your Own AI Agent Execution Trace in Python.

Wait 5 sec.

AI agents are getting better at doing things, not just answering questions. They can call APIs, search databases, execute functions, inspect files, make decisions, and continue working through a task without waiting for a human after every step, and that sounds great until something goes wrong.A normal application can usually give you a reasonably clear answer when a request fails. You can look at the HTTP request, inspect the application logs, check the database, and work backwards from the error. With an AI agent, the failure can be much harder to understand because the system is making decisions across multiple steps.The final answer might look perfectly reasonable even though something went wrong three or four steps earlier.An agent might have selected the wrong tool. A tool might have returned stale information. The model might have received an incomplete context. An API might have taken eight seconds to respond. A retry might have executed an operation twice. Or the agent might simply have decided to continue after receiving an unexpected result.When all you have is the final response, these failures are difficult to reconstruct, and that is why I think every serious AI agent needs something more fundamental than a beautiful chat interface: an execution trace.In this article, we are going to build a small execution-tracing system in Python from scratch. The goal is not to recreate a commercial observability platform. The goal is to understand what an agent actually needs to record while it is working, how that information can be structured, and how those traces can eventually become the foundation for debugging and monitoring production agents.The Problem With Looking Only at the Final AnswerImagine an agent that receives this request: “Find the latest order for customer 1842, check its payment status, and create a support ticket if the payment failed.”From the outside, this looks like one task, but from the agent's perspective, it might actually be a sequence of operations:Receive task ↓Identify customer ↓Query orders ↓Select latest order ↓Query payment service ↓Interpret payment response ↓Decide whether a ticket is required ↓Create support ticket ↓Return resultNow imagine the customer says that the agent created a ticket for the wrong order. If your system only stores the final response, you know that the agent created a ticket. You do not necessarily know why.Was the wrong order returned by the database?Did the agent select the wrong record?Did the payment API return an unexpected response?Did the model misunderstand the tool result?Did a retry happen?Did the customer have two orders created within a few seconds of each other?These are completely different failures. The first lesson is therefore simple: an agent's execution should be treated as a sequence of observable events, not as a single request and response and that changes how we think about logging. Instead of logging only:Agent completed task successfully.We want to know what happened throughout the execution.What Should an AI Agent Trace Contain?Before writing any Python, it helps to decide what we actually want to capture. For a simple agent, I would start with these fields:task_idagent_idstepmodeltoolargumentslatencyresulterrorretry_countfinal_statusEach field answers a different question. task_id tells us which execution the event belongs to. Without it, logs from concurrent agent executions can become almost impossible to follow. agent_id identifies the agent or workflow that produced the event. This becomes important when an application contains multiple agents with different responsibilities. step tells us where we are in the execution. An agent might perform ten or twenty operations during a single task, so step information gives us a basic timeline. model tells us which model participated in the decision. This becomes particularly useful when an application changes models or uses different models for planning, tool selection, and final responses. tool identifies an external operation. This could be a database query, web search, CRM API, payment service, internal function, or another agent. arguments Tell us what was actually sent to the tool. latency tells us how long the operation took. result records what came back. error captures failures. retry_count tells us whether the system attempted the operation more than once, and final_status gives us the overall outcome.The important thing here is that these fields are not just for developers staring at logs. Together, they form a story of the agent's execution.Let's Build a Small TracerWe will keep the first version deliberately simple. I don't want to hide the mechanics behind a framework. If you understand the basic implementation yourself, it becomes much easier to understand what larger observability systems are doing underneath. We'll start with a Python class that creates and stores trace events.import timeimport uuidfrom dataclasses import dataclass, fieldfrom typing import Any, Optional@dataclassclass TraceEvent: task_id: str agent_id: str step: int model: Optional[str] = None tool: Optional[str] = None arguments: Optional[dict] = None latency_ms: Optional[float] = None result: Any = None error: Optional[str] = None retry_count: int = 0 final_status: Optional[str] = Noneclass AgentTracer: def __init__(self, agent_id: str): self.agent_id = agent_id self.task_id = str(uuid.uuid4()) self.events = [] self.step = 0 def start_step(self): self.step += 1 return time.perf_counter() def record(self, **kwargs): event = TraceEvent( task_id=self.task_id, agent_id=self.agent_id, step=self.step, **kwargs ) self.events.append(event) def get_trace(self): return self.eventsThere is nothing particularly sophisticated here, and that is intentional. We are creating a trace identifier for an agent execution, maintaining a step counter, and storing structured events. Now we can execute a task and associate every operation with the same task_id. That single identifier becomes extremely valuable once several users are interacting with the application simultaneously.Why a Task ID Matters More Than It LooksSuppose your application has 500 users and the agents are processing requests concurrently. Your logs might contain:Agent startedTool calledTool completedAgent startedTool calledTool failedTool completedWithout a correlation identifier, those messages are ambiguous. Which tool call belonged to which user? A task_id change that.task_id=8f12step=1tool=customer_lookuptask_id=42acstep=1tool=customer_lookuptask_id=8f12step=2tool=order_lookupNow the execution can be reconstructed even when events from different tasks are interleaved. This is a pattern developers have been using in distributed systems for years, but AI agents make it even more important because a single user request can involve many models and tool interactions. I would therefore treat the task ID as the equivalent of a thread running through the entire agent execution.Recording Tool CallsThe next step is connecting the tracer to actual tool execution. Let's create a deliberately simple customer lookup function.def get_customer(customer_id): time.sleep(0.15) return { "customer_id": customer_id, "name": "Rahul", "plan": "enterprise" }We can wrap the call with our tracer.tracer = AgentTracer("support-agent")start = tracer.start_step()try: result = get_customer(1842) latency = (time.perf_counter() - start) * 1000 tracer.record( tool="get_customer", arguments={"customer_id": 1842}, latency_ms=latency, result=result, retry_count=0 )except Exception as exc: latency = (time.perf_counter() - start) * 1000 tracer.record( tool="get_customer", arguments={"customer_id": 1842}, latency_ms=latency, error=str(exc), retry_count=0 )Now we have something more useful than a text log. We know which task made the call, which step it was, what tool was used, what arguments were supplied, how long it took, and what came back, and that is already enough to answer several debugging questions.Measuring Latency Is Not OptionalOne thing I would strongly recommend recording from the beginning is latency. An agent can technically succeed while still being unusable. Imagine an agent that normally completes a task in three seconds. One day, the same task takes 45 seconds. The final response might still be correct.If you only record success or failure, your monitoring system will say everything is fine. But the user experience has deteriorated significantly. Latency allows us to identify where the time went. For example:Step 1 — retrieve customer 120 msStep 2 — retrieve order 180 msStep 3 — payment API 7,820 msStep 4 — model decision 920 msStep 5 — create ticket 210 msThe agent did not suddenly become “slow.” The payment dependency became slow. That distinction matters when debugging production systems.Recording Model DecisionsTool calls are only part of the execution. The model itself is another important component. Suppose the agent uses GPT-style reasoning to decide which tool to call. We should be able to see that a model decision occurred between two tool operations. A simplified event might look like this:start = tracer.start_step()model_name = "example-model"# In a real application this would be an API call.decision = { "tool": "get_customer_orders", "arguments": { "customer_id": 1842 }}latency = (time.perf_counter() - start) * 1000tracer.record( model=model_name, tool="get_customer_orders", arguments=decision["arguments"], latency_ms=latency, result=decision)In a production implementation, you would likely capture additional information such as token usage, model version, request metadata, and evaluation information. But even this simplified representation gives us a valuable timeline. We can now distinguish between: The tool failed, and The agent chose the wrong tool. Those are very different engineering problems.Errors Need ContextOne of the biggest mistakes in application logging is recording errors without enough surrounding information. Consider:ERROR: Payment API failedThat tells us almost nothing. A better trace might say:task_id=8f12step=4tool=payment_statusarguments={"order_id":"ORD-9281"}latency_ms=8120retry_count=2error="HTTP 504 Gateway Timeout"Now we know what happened. The request was made for an order ORD-9281It took more than eight seconds, the system retried twice, and the dependency ultimately timed out. That information can lead us toward an actual fix. It may also reveal another problem: perhaps retrying a particular operation is unsafe. This is where execution traces start becoming more than debugging logs. They become a way to understand the behaviour of the entire system.Retries Can Quietly Create Dangerous BugsRetries are useful, but with agents, they can introduce another class of problems. Imagine an agent is instructed to create a support ticket. The request reaches the ticketing service, and the ticket is successfully created.But the response gets lost because of a network timeout. The agent sees a timeout and retries. Now there are two tickets. From the agent's perspective, the first attempt looked like a failure and from the external system's perspective, the operation succeeded. A trace helps expose this situation.step=7tool=create_ticketretry_count=0result=timeoutstep=7tool=create_ticketretry_count=1result=successAt first glance, this looks normal. But if we inspect the external system and find two tickets, we have discovered a classic distributed-systems problem: the client did not know whether the first operation had completed.This is one reason I prefer recording retry information as part of the trace rather than treating retries as invisible implementation details.Adding a Retry WrapperWe can build a simple retry mechanism around our tools.def execute_with_retry( tracer, tool_name, function, arguments, max_retries=2): for attempt in range(max_retries + 1): start = time.perf_counter() try: result = function(**arguments) latency = (time.perf_counter() - start) * 1000 tracer.record( tool=tool_name, arguments=arguments, latency_ms=latency, result=result, retry_count=attempt ) return result except Exception as exc: latency = (time.perf_counter() - start) * 1000 tracer.record( tool=tool_name, arguments=arguments, latency_ms=latency, error=str(exc), retry_count=attempt ) if attempt == max_retries: raiseThis is still a small example, but now the retry behaviour is visible. In a real application, I would not blindly retry every exception. A timeout may be retryable. An authentication error probably is not. A malformed request should generally not be retried without changing something. The trace gives us the evidence needed to make those decisions intelligently.Turning the Trace Into a TimelineOnce events are stored, we can print them as a human-readable execution timeline.def print_trace(events): print(f"\nTask: {events[0].task_id}") print(f"Agent: {events[0].agent_id}") print("-" * 70) for event in events: status = "ERROR" if event.error else "OK" print( f"Step {event.step} | " f"{event.tool or 'MODEL'} | " f"{status} | " f"{event.latency_ms:.2f} ms" )The output could look something like:Task: 8f12...Agent: support-agent----------------------------------------------------------------------Step 1 | get_customer | OK | 153.22 msStep 2 | get_orders | OK | 204.19 msStep 3 | payment_status | ERROR | 8012.44 msStep 3 | payment_status | OK | 311.82 msStep 4 | create_ticket | OK | 188.51 msNow the agent execution is no longer a black box. We can see the sequence, failure, retry and how long each operation took. And because every event belongs to the same task, we can reconstruct what happened.From Logs to a Small DashboardOnce you have structured events, creating a basic dashboard becomes surprisingly straightforward. For a first version, I would display four things:Total tasks: This tells us how much work the agent has processed.Success rate: This shows how many tasks were completed without an unrecovered error.Average latency: This provides a high-level view of performance.Tool failures: This shows which dependencies are creating problems.A simple Python application could expose the data through a small API and render it using a lightweight frontend. For example, you could transform the trace into JSON:import jsonfrom dataclasses import asdicttrace_json = json.dumps( [asdict(event) for event in tracer.get_trace()], indent=2, default=str)print(trace_json)That JSON can then be consumed by a dashboard. The important point is that the dashboard isn't the foundation. The trace is the foundation. Once the underlying execution data is structured properly, you can decide later whether you want a terminal view, a web dashboard, a database, OpenTelemetry integration, or another observability system.What a Real Agent Trace Starts to Look LikeWith several steps, the trace begins to resemble a distributed execution graph rather than a simple log. For example:Task: 8f1201 MODEL Decide customer lookup02 TOOL get_customer 153 ms03 MODEL Decide order lookup04 TOOL get_orders 204 ms05 MODEL Decide payment lookup06 TOOL payment_status 8012 ms timeout07 TOOL payment_status retry=1 312 ms08 MODEL Decide ticket creation09 TOOL create_ticket 189 ms10 MODEL Generate final responseThis is much closer to how I think about an agent internally. It isn't really “a chatbot.” It is an execution system that happens to use a language model as one of its decision-making components. Once you see it that way, tracing becomes much more important.Don't Log Everything BlindlyThere is an important warning here. More logs do not automatically mean better observability. Agent traces can contain extremely sensitive information. Tool arguments might contain customer data. Model context might contain internal documents. API responses might contain credentials or personal information. If you simply dump every prompt, response, and tool result into a database, you can create a security problem while trying to solve an observability problem.I would therefore separate debug information from sensitive application data. For example, instead of storing an entire customer record, the trace might store:customer_lookupcustomer_id_hashrecord_found=truelatency=142msSimilarly, instead of storing an authentication token inside tool arguments, the trace should record that authentication was used without exposing the credential itself. The trace should help engineers understand execution without becoming a second copy of the application's most sensitive database.What Should We Do With Model Prompts?This is another area where I think developers need to be careful. During development, having the exact prompt and model output can be extremely useful. In production, however, storing every prompt and response indefinitely may not be appropriate. A better architecture can separate: Trace metadata from Content payloads. The trace can record:task_idstepmodellatencytoken_counttoolstatuswhile the actual prompt or response is stored separately under controlled retention policies, if it needs to be retained at all. This separation also makes the tracing system easier to scale.A Trace Should Tell You Why, Not Just WhatThis is probably the most important idea in the entire implementation. Traditional logs often answer: What happened?An agent trace should help answer: What happened, in what order, under which context, and where did the execution diverge from what we expected?Suppose an agent makes an incorrect decision. The final output alone tells us:Agent chose option B.The trace may tell us:Step 4:Model received context version 17.Step 5:Database returned record version 18.Step 6:Agent used stale context.Step 7:Tool was called with outdated information.That is a completely different level of debugging. We are no longer guessing about the agent. We are examining its execution.Extending the TracerThe simple implementation we built is only the beginning. If I were extending this into a production-oriented tracing layer, I would add several more concepts.The first would be token usage. Tracking input and output tokens can help identify unexpectedly expensive workflows.The second would be the model version. If behaviour changes after a model update, the trace should allow us to identify which executions used which version.The third would be parent-child relationships. A tool call could become a child span of a model decision, and an entire agent execution could become the parent trace.The fourth would be timestamps. Duration is useful, but absolute timestamps make it easier to correlate agent events with database logs, API gateway logs, infrastructure events, and other systems.The fifth would be status categories rather than a simple success/failure flag. For example, an operation could be successful, retryable, rejected, cancelled, timed out, or awaiting human approval.The sixth would be evaluation metadata. After an agent completes a task, we may want to record whether the output was later judged correct, incorrect, safe, or useful.That last part is particularly interesting because it connects observability with evaluation.The Trace Can Become Your Dataset for Improving the AgentOnce an agent has been running for a while, its execution traces become extremely valuable. Suppose you discover that 8% of tasks fail because a particular tool returns an unexpected response. You now have real production examples. You can take those failures and turn them into regression tests. Instead of simply saying, “We should make the agent more reliable,” you can create a test suite containing actual failure patterns. For example:Test 001Tool returns empty response.Test 002Tool times out.Test 003Tool returns malformed JSON.Test 004Database record changes between reads.Test 005Agent attempts unauthorized operation.The tracing system, therefore, becomes part of the improvement loop. The agent executes. The system records what happened. Failures are identified. Those failures become tests. The tests are used to improve the agent. The improved agent generates new traces. Over time, the system becomes easier to reason about because you are learning from actual executions rather than only from synthetic examples.Where OpenTelemetry FitsOnce the basic concept is understood, you don't necessarily need to maintain a custom tracing system forever. OpenTelemetry provides a standard approach for collecting telemetry across distributed applications. That becomes particularly interesting for AI agents because an agent is often a distributed system in disguise. A single task may involve:Application ↓Agent ↓Model API ↓Database ↓Internal API ↓Third-party serviceIf all of these components emit compatible telemetry, engineers can follow a request across system boundaries. The custom Python tracer we built in this article is therefore not intended to replace mature observability infrastructure. Its purpose is to make the underlying idea understandable. Before reaching for a large framework, it is useful to understand what information you actually need.From Agent Logs to Agent ObservabilityThere is a subtle difference between logging and observability. Logging tells you what the application decided to write down. Observability is about having enough information to understand the internal state of a system from its external outputs. For traditional applications, metrics, logs, and traces have become standard tools for achieving this. AI agents add another dimension because the system contains probabilistic decision-making. The same high-level task can sometimes produce different execution paths. That means we need to observe not just whether the request succeeded, but how the agent arrived there. A trace gives us that execution history.It lets us ask questions such as:Which tools are failing most often?Which steps consume the most time?Which model calls are generating expensive workflows?Where are retries occurring?Which agent tasks require human intervention?Which failures repeat across customers?Which tool responses cause unexpected behaviour?Which model versions are associated with particular failure patterns?Those questions are much harder to answer from a simple application log.The Bigger LessonWhen developers first build AI agents, it is natural to focus on the model. We think about which model to use, how to improve the prompt, how much context to provide, and how to make the agent reason better. Those things matter.But once the agent starts interacting with real systems, another reality appears. The model is only one part of the execution. The agent depends on APIs, databases, authentication, network connections, tool definitions, application state, retries, permissions, external services, and data quality. Any one of those components can influence the final result, and that is why I think an execution trace should be considered part of the agent architecture rather than something added after the first production incident.If an agent can make ten decisions and call five different tools, we should be able to reconstruct those ten decisions and five tool calls when something goes wrong. Otherwise, we are effectively asking engineers to debug a distributed system by looking only at its final sentence.ConclusionBuilding an AI agent is relatively easy compared with understanding what that agent does once it starts operating inside a real application.The first version of an execution tracer does not need to be complicated. A task identifier, agent identifier, step number, model information, tool name, arguments, latency, result, error, retry count, and final status already give us a much clearer picture of the system. From there, the architecture can grow. You can add timestamps, token usage, model versions, parent-child spans, evaluation results, security controls, dashboards, OpenTelemetry, and long-term analytics. But the fundamental idea remains the same.An AI agent should leave behind a useful record of how it worked. Not because developers want to watch every decision an agent makes, but because production systems eventually fail in ways that are impossible to understand from the final answer alone.When that happens, the question is no longer simply: “What did the agent say?”The more useful question is: “What happened during the execution that caused the agent to say it?”That is what an execution trace gives us, and for developers building agents that are expected to operate reliably outside a demo environment, that distinction can make the difference between debugging by guesswork and debugging from evidence.