Observability has long been a staple of enterprise operations, giving companies a way to understand what their applications and infrastructure are doing, where they are failing, and why.The arrival of LLMs and AI agents, however, complicates matters. An application can be running perfectly well from an infrastructure perspective while the agent sitting on top of it gives the wrong answer, calls the wrong tool, or fails to complete the task it was given.And this is precisely why observability software company Dynatrace announced in mid-August its intent to acquire Arize for a cool $915 million. The idea, ultimately, is to bring Dynatrace’s existing visibility into applications and infrastructure together with Arize’s ability to trace, evaluate and debug the behavior of models and AI agents.That deal has now officially closed, and The New Stack caught up with Dynatrace CPO Steve Tack and Arize co-founder and CPO Aparna Dhinakaran to dig into what bringing the two platforms together actually means — and why the arrival of AI models and agents “demands a new kind of observability.”Two sides of observabilityDynatrace, for the uninitiated, has more than two decades of history in application performance monitoring (APM), a remit that has since expanded into full-stack observability and security. Founded out of Austria in 2005, the company was acquired by Compuware for $256 million in 2011, with private equity giant Thoma Bravo in turn acquiring Compuware in 2014. The following year, Thoma Bravo carved Dynatrace out as a standalone business, merged it with fellow portfolio company Keynote Systems, and eventually took Dynatrace public on the New York Stock Exchange in 2019.Tack has been there for much of that journey, having spent more than a decade at Compuware, before joining Dynatrace in 2012 as senior vice president of product management, followed by chief product officer from 2024.AI has become an increasingly important part of that evolution. In 2017, Dynatrace formally launched Davis, its AI-powered digital assistant for performance monitoring, which subsequently evolved into the company’s causal AI engine for identifying the root causes of problems. The company has added predictive and generative AI capabilities and, more recently, agentic AI including autonomous SRE agents designed to investigate and remediate incidents.“The one thing that we’ve done on the Dynatrace side for a long time is to invest in the use of AI to power observability.”“The one thing that we’ve done on the Dynatrace side for a long time is to invest in the use of AI to power observability — that’s in our DNA,” Tack tells The New Stack.That history also gives Dynatrace common ground with Arize, Tack says, particularly around the latter’s work with Signal, an agent that reviews production traces, identifies recurring problems and surfaces likely causes and potential fixes.Signal by ArizeArize, for its part, emerged from stealth in 2020 as an AI observability startup focused on helping companies monitor and troubleshoot machine learning models in production. As LLMs and agents have become more prevalent, it has expanded into tracing agent behavior, running evaluations and identifying failures that conventional application monitoring might never register.And that gets to the crux of why the two companies see themselves as complementary. Dynatrace has traditionally focused on the applications, services and infrastructure beneath a system, serving SRE and platform engineering teams; Arize has focused more on the behavior and output of AI models and agents, with AI engineers and developers as its core audience.But those layers increasingly overlap. Dhinakaran notes that an AI agent typically depends on a host of conventional software components and APIs to actually carry out a task. Arize might be able to identify a problem with an agent’s reasoning, tool use or output, but if the failure originates in one of the underlying services it calls, developers still need visibility into the software beneath it.“I think about it like this — we used to have one side of the coin, the ability to debug all the harness and the LLM-related issues, but if it came to a software issue, we have to then go look at our software traces to figure out the root cause,” Dhinakaran tells The New Stack. “Now, we can put up way more improvements, because we have both sides of the coin to be able to debug.”And so bringing the two platforms together means an agent failure can be investigated across both layers. In a statement provided to The New Stack, Stephen Elliot, IDC group VP for software development and IT operations, says that as agents proliferate across the enterprise, the need to align two separate but closely related parts of the AI application stack is more pressing than ever.“Bringing evaluation and observability together closes the loop between building AI applications and running them reliably in production, giving teams the visibility they need to catch issues earlier and resolve them faster,” Elliot says.Taking actionFor Dhinakaran, the bigger shift perhaps lies in who — or what — consumes all the juicy observability data. As software systems generate ever more telemetry and become increasingly autonomous, she sees observability moving beyond humans manually inspecting dashboards and traces, toward agents that can interpret that data and act on what they find.“It’s cool that observability is no longer about just looking at data. It’s actually about going from the data to taking action.”“The observability category reinvents itself every couple of years,” Dhinakaran says. “As the industry changes and new tools come out, this category has always had to adapt. And right now, it’s cool that observability is no longer about just looking at data. It’s actually about going from the data to taking action.”Part of that shift is simply a matter of volume. Modern applications can generate far more telemetry than engineers could reasonably inspect themselves, creating an obvious role for agents that can continuously sift through it on their behalf.“No human wants to go look at billions of traces,” Dhinakaran continues. “Nobody’s going to go do that. And so how do you have agents go read your telemetry data?”“No human wants to go look at billions of traces.”There are already signs of what this might look like in the real world. Anthropic recently disclosed a series of cybersecurity evaluation incidents in which Claude models gained access to real third-party systems. In investigating them, Anthropic used agentic search to sift through huge volumes of model transcripts, with Claude itself helping review millions of conversations flagged for closer inspection.Dhinakaran sees that same principle extending into everyday software operations: once an agent can understand the telemetry, the next logical step is allowing it to do something with what it finds. Or “take action”, in other words.“I think that agents are going to be the primary consumers of a lot of this data, and then because they’re going to consume that data, if they have access to repos, if they have access to other skills, well, now they can go put up a fix,” Dhinakaran says.“I think that agents are going to be the primary consumers of a lot of this data.”Arize already does this with Alyx, an assistant inside its own product. Signal reviews Alyx’s traces, spots recurring failures with the same root cause, and opens pull requests to fix them, around “65% to 70%” of which Arize accepts, according to Dhinakaran. So the engineer who owns Alyx has gone from combing through traces, to reviewing PRs.Dhinakaran sees that as a sign of where things are heading. “Agents review all the traces — agents become the first responders, and agents put up pull-requests,” she says.The longer-term idea, as Dhinakaran puts it, is “self-sustaining, self-maintaining software, but also kind of self-improving,” with humans still doing the final review.Buy vs build, and what happens nextIt’s worth noting that the two companies weren’t formal partners before the acquisition, though they did have mutual customers — and Tack says some of those customers had already suggested they would like to see Dynatrace and Arize operating more like a single entity. Part of the rationale is that AI agents don’t operate in isolation: they still depend on the applications, services, cloud platforms and infrastructure around them, which makes the boundary between AI observability and conventional software observability increasingly difficult to separate.“No one’s got to convince anyone that there’s a massive shift happening from coding agents and from autonomous work,” Tack says. “But all these systems are also living in a broader software ecosystem.”That overlap also feeds into the familiar build vs buy quandary many companies find themselves in. Tack points to some of the assets Arize had already developed as part of the “buy” attraction on Dynatrace’s side. This includes Phoenix, a source-available platform for tracing, evaluating and debugging AI applications and agents; and OpenInference, which builds on OpenTelemetry with AI-specific conventions for capturing things such as LLM calls, tool-use, retrieval, and agent behavior.“Those are not easy things to build,” Tack says.There was also the question of whether a conventional partnership would have been enough. Tack says some degree of integration might have been possible that way, but Dynatrace ultimately saw a deeper combination as necessary.“Some of it could maybe have been done through partnership, but I don’t think the level of the seamless experience on the end-to-end side would be achievable through a partnership,” Tack says.On a more practical level, that doesn’t mean Arize disappears into Dynatrace overnight. Arize will continue supporting both Phoenix and AX, its enterprise platform, while Arize capabilities are integrated into Dynatrace over time.Dhinakaran similarly says that Arize will remain available as a standalone product, meaning customers won’t need to be Dynatrace users to adopt it. But the longer-term goal, ultimately, is to connect two stages that have typically been treated separately: evaluating AI applications as they are being built, and understanding how they behave once they reach production.The post “No human wants to look at billions of traces”: Dynatrace bought Arize because agents need a new kind of observability appeared first on The New Stack.