Datadog launches AI-powered tools to bridge gap between system health and customer success

A system can be running perfectly and still fail its users.
The opening observation that sets up the problem Datadog's new tools are designed to address.
Mark

So the core problem here is that a company's systems can look fine on every metric and still be failing customers. How common is that actually?

Mimi

It's become pretty common, especially as software changes faster. You deploy code, monitoring says everything is green, but customers can't complete checkout. The systems are up. The latency is good. But something in the flow is broken.

Luke

The source doesn't give us numbers on how often this happens or how much revenue it costs companies. We know it's a problem Datadog is trying to solve, but we don't have data on the scale.

Mark

Fair point. So Journey Monitoring—it's pulling together data that already exists in Datadog's platform, right? It's not measuring something new.

Mimi

Exactly. It's combining Real User Monitoring, which watches actual customers, with Synthetic tests and Product Analytics. The novelty is putting them in one view organized around user journeys instead of keeping them separate.

Luke

And the claim is that this helps teams spot conversion drops faster. But the source doesn't show us a case study or a before-and-after. We're taking Datadog's word that this actually reduces the time to detect problems.

Mark

What about Bits Testing? That sounds like it's trying to solve a different problem—keeping tests from breaking when you redesign the interface.

Mimi

Right. Instead of recording a script that says "click button at coordinates 100, 200," it tests for the outcome. Did the user complete the task? That way, if the button moves, the test still works.

Luke

The AI-generated test suites from plain-language prompts—that's interesting, but the source doesn't tell us how accurate those generated tests are or whether they miss edge cases. We know Datadog claims it works, but there's no independent validation here.

Mark

So these are both real problems, but we're mostly seeing Datadog's framing of the solution.

Mimi

Yes. And the framing makes sense—the gap between system health and customer success is real. But whether these specific tools actually close that gap at scale, we'd need to see in practice.

Luke

The other thing worth noting: this is a company selling more tools to its existing customers. That's not a criticism, but it's context. Datadog benefits if teams feel like they need more visibility and more testing.

  • Engineering teams have long faced a quiet crisis: servers can run perfectly while customers silently fail to check out, log in, or finish onboarding — and no existing dashboard shows the full picture.
  • Datadog's Journey Monitoring pulls together conversion rates, live traffic patterns, and availability signals into a single view, surfacing the routes users actually take and flagging when completion rates begin to fall.
  • Bits Testing attacks the problem from the other direction — generating full test suites from plain-language prompts and re-discovering user paths at runtime, so tests don't shatter every time a designer moves a button.
  • The outcome-based approach is especially urgent for AI-powered features, where outputs vary by nature and scripted tests that expect identical responses become instantly obsolete.
  • Datadog's Chief Product Officer frames both tools as AI-as-infrastructure — not a feature layer, but a continuous engine for discovering what matters and verifying that it keeps working as products evolve.

A system can be technically flawless and still leave its users stranded — a paradox that has quietly grown into one of the defining challenges of modern software. Datadog, a widely used observability platform, this week introduced two tools designed to bridge the distance between machine health and human success: Journey Monitoring, which unifies conversion, traffic, and availability data around the paths users actually take, and Bits Testing, which uses AI to generate adaptive tests that follow outcomes rather than rigid scripts. The launch reflects a broader reckoning in the industry — that as software grows more complex and AI-driven, the old measures of what is 'working' are no longer enough.

A system can be running perfectly and still fail its users. Servers respond in milliseconds, code deploys without errors, infrastructure holds steady — and yet customers abandon their carts mid-checkout or never finish onboarding. This gap between what machines report and what people actually experience has become one of the harder problems in modern software operations.

Datadog moved this week to close that gap with two new tools: Journey Monitoring and Bits Testing. Rather than asking only whether a system is up, teams can now ask whether customers are actually succeeding at the things that matter — completing a purchase, resetting a password, finishing account setup.

Journey Monitoring pulls together three streams of data that typically live in separate places: conversion metrics, traffic patterns, and availability signals, organized around specific user journeys. By layering Real User Monitoring, Synthetic Monitoring, and Product Analytics into a single view, engineering and product teams can see the same journey three ways at once — how the system performed, what customers actually did, and whether they reached their goal. Color-coded health indicators and service level objective badges flag when completion rates drop, and the tool maps how problems in one flow can cascade into related parts of the application.

Bits Testing works from the opposite direction, keeping tests relevant before code ships. Teams describe what they want to test in plain language, and the tool generates browser, API, network, and goal-based checks without requiring hand-written scripts. Crucially, it tests for outcomes rather than replaying fixed sequences — re-discovering paths at runtime so tests adapt as the product changes, rather than breaking the moment a designer repositions a button.

This distinction is especially significant for AI-powered features, where outputs vary from run to run and scripted tests expecting identical responses become useless. An outcome-based test that checks whether a user could complete their task despite that variation stays relevant.

Datadog's Chief Product Officer Yanbing Li described both tools as part of a philosophy in which AI functions as infrastructure running beneath everything — monitoring systems, securing networks, and now tracking whether customers succeed. As applications grow more complex, deploy more frequently, and incorporate AI-driven features as standard, the old ways of testing and monitoring have started to break down. Datadog's answer is to automate the discovery of what matters and the continuous testing of whether it works.

A system can be running perfectly and still fail its users. The servers respond in milliseconds. The code deployed without errors. The infrastructure holds steady. And yet customers abandon their shopping carts mid-checkout, or get stuck on a login screen, or never finish onboarding. This gap—between what the machines say is working and what actually happens when a person tries to use the software—has become one of the harder problems in modern software operations.

Datadog, the observability platform used by thousands of engineering teams, is moving to close that gap with two new tools launched this week: Journey Monitoring and Bits Testing. Together, they represent a shift in how companies think about software health. Instead of asking only whether the system is up, teams can now ask whether customers are actually succeeding at the things that matter—completing a purchase, resetting a password, finishing an account setup.

Journey Monitoring works by pulling together three streams of data that usually live in separate places. It combines conversion metrics (did the user finish?), traffic patterns (how many people tried?), and availability signals (was the system responsive?) into a single view organized around specific user journeys. The tool draws on Datadog's existing Real User Monitoring, which watches actual customers in production, alongside Synthetic Monitoring, which runs automated tests, and Product Analytics, which tracks how people actually use features. By layering these together, engineering and product teams can see the same journey three ways at once: how the system performed, what customers actually did, and whether they reached their goal.

The tool also watches live traffic to find the paths customers most often take to reach the same destination. If most people who successfully check out follow one route through the application while others get stuck on a different path, Journey Monitoring surfaces that pattern before it becomes a support ticket or an outage report. The interface uses color-coded health indicators and service level objective badges to flag when conversion rates drop, and it maps connections between journeys so teams can see whether a problem in one flow is cascading into related parts of the application.

Bits Testing tackles the problem from the opposite direction. Instead of monitoring what happens in production, it focuses on keeping tests relevant before code ships. The tool generates test suites from plain-language prompts—a team member can describe what they want to test in ordinary English, and Bits Testing creates browser tests, API tests, network tests, and goal-based checks without requiring anyone to write code for each one. The key innovation is that it tests for outcomes rather than replaying a fixed sequence of steps. A traditional scripted test fails the moment a designer changes a button's position, even if users can still click it and complete the task. Bits Testing re-discovers the paths at runtime, allowing tests to adapt as the product evolves.

This distinction matters especially for applications that incorporate artificial intelligence. When an AI model generates text or makes a decision, the output varies from one run to the next. A scripted test that expects an identical response every time becomes useless. An outcome-based test that checks whether the AI produced a reasonable result, or whether the user could complete their task despite the variation, stays relevant.

Datadog's Chief Product Officer Yanbing Li framed both tools as part of a larger philosophy about how AI should work in software operations. Rather than treating AI as a feature to bolt onto a product, the company sees it as infrastructure that should run underneath everything—monitoring systems, securing networks, and now tracking whether customers succeed. Journey Monitoring autonomously discovers the critical paths users take and observes them continuously. Bits Testing acts as an agent that keeps testing those paths as the product changes. Together, they represent what Li called a two-way approach: visibility and control on one side, continuous adaptation on the other.

The launch reflects a real tension in modern software development. As applications grow more complex, as teams deploy updates constantly, and as AI-driven features become standard, the old ways of testing and monitoring have started to break down. A service can be technically healthy while users struggle. Tests written for yesterday's interface fail on today's redesign. The gap between system metrics and customer success has widened, and it has become harder to close without new tools. Datadog's answer is to automate the discovery of what matters and the continuous testing of whether it works.

We don't think of AI as a feature you add to a product. We think of it as what should be running underneath everything—infrastructure, security and now the user journey.
— Yanbing Li, Chief Product Officer at Datadog
Quieres la nota completa? Lee el original en itbrief.com.au ↗
Contáctanos FAQ