AWS CloudWatch Omni Tackles the Core Problem of Agent AI: Explainability
Amazon's new observability platform shifts focus from system health to agent decision-making, addressing the gap between technical metrics and business outcomes that has kept agentic AI in pilot mode.
For the past several decades, observability tools have centered on a single question: Is the system operational? Agentic artificial intelligence upends this framework. An agent might deliver a polished response, hit its performance targets and generate no errors, yet still provide an incorrect answer to a user, invoke the wrong capability or reference outdated information. Traditional metrics show everything working fine, but the business sees failure. This disconnect is what Amazon Web Services Inc. is addressing with Amazon CloudWatch Omni, which reached general availability last week. AWS positions Omni as the successor to its existing CloudWatch offering.
The platform operates independently of the AWS Management Console, leverages OpenTelemetry standards and incorporates AI capabilities. This repositions observability away from "Is it running?" toward "Why did my agent do that?" That shift represents the critical hurdle separating most organizations from deploying agentic AI at scale. AWS references an IDC projection estimating over 1 billion deployed agents by 2029. No operations team can feasibly examine that volume of unpredictable behavior manually.
Evaluation is the new monitoring
The evaluation engine, rather than the dashboards, constitutes Omni's core strength. The platform records every trace and includes 17 pre-built evaluators measuring coherence, usefulness, accuracy and routing precision, among other dimensions. Users can test different prompts side-by-side in a sandbox environment, construct test cases from actual production data, and identify performance degradation automatically. Evaluators operate continuously on active traffic, flagging quality deterioration the way a traditional system would alert on resource consumption spikes. While latency and error counts cannot determine correctness, evaluation scoring can.
Sony has adopted the platform early. "At Sony, our enterprise-wide agentic AI platform now supports hundreds of proof-of-concept and production workloads," said Masahiro Oba, senior general manager of the AI Acceleration Division at Sony. "At this scale, observability and evaluation are essential. With Amazon CloudWatch Omni, I can go from a single trace directly to evaluation, AI analysis, comparison or dataset creation."
The word "hundreds" carries significance. Most organizations do not struggle with deploying a single agent; they face challenges managing dozens or hundreds, each developed by separate teams with distinct quality standards. Oba highlighted that creating evaluation datasets frequently becomes a business constraint, and one-click dataset generation from production traces eliminates this friction.
Getting out of the console is a bigger deal than it sounds
Developers receive a native plugin for Visual Studio Code, Cursor and Kiro, displaying traces as agents execute locally without needing an AWS account. Operations personnel access a standalone interface with authentication through existing providers like Okta and Microsoft Entra ID. Both environments connect to the same underlying data, meaning the trace a developer examines matches the one an operator reviews.
AWS recognizes that its console targets infrastructure administrators rather than site reliability engineers, AI engineers and application owners who now shoulder operational duties. Bringing tools into the development environment, where AI assistants such as Claude Code and Codex can configure monitoring, moves quality assurance earlier in the development cycle where corrections cost less.
The unified data layer is the real differentiator
Numerous startups can instrument language model interactions. Omni's distinction lies in consolidating agent traces, application metrics and infrastructure data into a single CloudWatch repository. An investigation can therefore begin with an agent receiving faulty tool output, progress to an application programming interface failure from a resource-constrained backend, and conclude with exhausted database connections.
In typical environments today, this scenario requires three separate platforms, three teams and extensive coordination to piece together the root cause. AWS DevOps Agent activates automatically during investigations, linking signals and preserving a complete investigation record.
Capital One participated in the design phase. "Capital One operates one of the largest observability footprints in financial services," said Parvez Naqvi, managing vice president of cloud platform and resilience engineering at Capital One. "As a design partner for Amazon CloudWatch Omni, we helped shape a single AI-powered observability solution that will give our engineers topology-aware intelligence and natural-language querying across all telemetry from a single surface, with full data ownership through OpenTelemetry."
For regulated sectors like banking, data portability proves essential, which demands data ownership. Preserved investigation records also matter significantly, as they constitute audit documentation of how an AI incident was managed in regulated environments.
Open standards lower the lock-in bar but don't eliminate it
Instrumentation relies on OpenInference and the AWS Distro for OpenTelemetry, regardless of whether agents operate on AWS or alternative cloud providers. Omni accommodates LangChain, LangGraph, CrewAI, the OpenAI Agents SDK, Strands and the Vercel AI SDK, plus external evaluators including DeepEval. Amazon Bedrock AgentCore agents gain Omni capabilities automatically. Azure support exists today, with expanded multicloud integration planned.
This interoperability is necessary given intense competition. Datadog, Dynatrace, New Relic, Grafana Labs and Splunk are expanding agent observability offerings, alongside language model-specific platforms such as LangSmith and Arize AI.
OpenTelemetry ensures data portability, but the analytical layer does not. Topology mapping, evaluators, investigation records and DevOps Agent all execute on AWS infrastructure. Organizations with substantial AWS commitments will likely adopt Omni as their default. Those maintaining mature Datadog or Splunk deployments spanning multiple clouds may use Omni for agent prototyping and testing while retaining their primary observability system elsewhere.
Pricing is built for adoption, so watch the telemetry bill
The IDE extension carries no cost. Customers pay for transmitted and retained telemetry. Dashboards and alerts incur no charges, and query allowances extend to five times monthly ingestion volume. Qualifying accounts receive a 30-day trial period and $1,000 in OpenTelemetry ingestion credits.
The challenge is that agents generate substantial data. Every request, model invocation, tool call and agent-to-agent transfer creates a span. When multiplied across hundreds of deployments plus continuous evaluation, ingestion expenses can exceed the AI budget that funded them. DevOps Agent incurs separate charges.
What this means for buyers
Omni represents a full-featured agent observability solution bridging the confidence gap preventing agents from advancing beyond experimental phases. Technology and business leaders should consider:
- Establish success criteria before implementing the evaluator. Integrated scoring only delivers value when business stakeholders have defined what constitutes a correct, compliant and valuable response for each agent. Few organizations have completed this work.
- Standardize OpenTelemetry instrumentation across all agent experiments now. This preserves flexibility and streamlines future technology decisions.
- Calculate telemetry expenses at full production volume. Establish guidelines for data sampling, storage duration and evaluation frequency before agents reach production, not after receiving the initial bill.
- Clarify where your primary observability system operates. AWS as your main cloud makes Omni a logical choice. For multicloud environments with established observability infrastructure, consider Omni for development and evaluation, then integrate rather than replace existing systems.
- Incorporate investigation records into governance frameworks. Integrate Omni's investigation documentation into AI governance and compliance programs, particularly for regulated industries.
The industry invested a decade mastering distributed system observation. Agentic AI demands observing decisions rather than infrastructure, and AWS intends to control that layer for its customer base. Organizations maximizing value will treat evaluation as a core operational practice, not merely an optional capability.


