← Back to blog
AiAbout 6 min read

ServiceNow Turns Agent Failures Into Training Data

Published Oct 4, 2026
ServiceNow Turns Agent Failures Into Training Data

Every team that has deployed an AI agent into a real enterprise system knows the same pain. The model is broadly capable, and then it meets your specific tools, your specific policies, your specific mess of a state machine, and it stumbles.

ServiceNow's research arm, CoreAI, has released a system that tries to turn those stumbles into an asset. It is called AutoSynthData, and its premise is that a model's failures are the most valuable raw material you have for training it.

The central move

The pipeline starts by running a target model and a stronger teacher model against the same diagnostic tasks in a real environment. Where the target model fails and the teacher succeeds, there is a capability gap. AutoSynthData distills that gap into what the team calls capability specification cards: sanitized descriptions of the tools, workflows, end states and permissible variations involved, stripped of the original evaluation prompts and entity details.

Those cards seed the generation of new tasks. Crucially, the generator never sees the original evaluation trajectories, so the resulting data tests the underlying competence rather than memorized sequences. The framework scales this in two phases. A target phase creates core tasks in parallel. A multiply phase takes the verified tasks and spins out variants with new wording, new entity configurations and new initial states. One rule keeps the data from drifting: a multiplied sample can never seed another multiplication.

Three conditions for a task worth generating

ServiceNow is explicit that a plausible-looking request is not enough. A generated task must satisfy three things at once.

It has to be feasible, meaning there is at least one executable path through the environment that completes the request while respecting the rules. It has to be realistic, meaning it mirrors work a real user would actually ask for, rather than a technically valid action nobody would take. And it has to be difficult, meaning it targets a genuine weakness. A trivial task teaches nothing, and an impossible one teaches the wrong lesson.

Each task ships with a verifier, and the verifier has to meet its own bar. It must be consistent with the prompt and system state, sound enough to reject constraint violations, and complete enough to accept any working solution rather than insisting on one reference path.

The quality gates are the real product

The part of this that deserves attention is not the task generation. It is the checking.

Every candidate passes two gates. A positive gate runs the reference solution inside the environment to confirm the intended answer actually satisfies the verifier, which catches mismatches between the prompt, the initial state and the success criteria. A negative gate then deliberately mutates the final state to confirm that wrong outcomes are actually rejected. That second gate is the one that catches under-specified verifiers, the kind that would happily reward a malformed agent trajectory.

When a candidate fails, a critic diagnoses the breakdown, spotting broken references or contradictory states, and directs a bounded repair rather than discarding the work outright. Above the sample level, a batch review watches the aggregate: which task clusters are overrepresented, which dimensions are missing, which generation patterns keep stalling. The controller then steers the next batch toward the gaps that remain.

Why anyone needs synthetic data at all

It is worth stepping back to ask why this matters. The ideal training data for an enterprise agent is a record of real people doing real work in the real system, correctly labeled. That data is scarce, sensitive and expensive to collect, and much of it cannot leave the building for privacy or compliance reasons.

Synthetic generation is the escape hatch, but it comes with a known failure mode. If you generate tasks from a model, you inherit that model's blind spots, and the resulting data can look plentiful while teaching very little. The whole contribution of AutoSynthData is the machinery around the generation that keeps the data honest: gates that reject bad tasks, diversity checks that prevent repetition, and a curriculum that keeps aiming at what the agent has not yet mastered. Generation without verification is noise. This is an attempt to make it signal.

What the numbers say, and do not say

In tests on EnterpriseOps Gym, an open-source environment for enterprise agent tasks, a Gemma-based model fine-tuned on 2,000 AutoSynthData samples improved its Pass@1 metric by 35 percent relative to baseline in a hybrid domain. In the ITSM domain, synthetic supervised fine-tuning lifted mean Pass@1 from 18.77 percent to 27.18 percent.

Those are real gains on a specific benchmark, which is not the same as a general result. The pipeline's central dependency is honest and worth stating plainly: it needs a significantly stronger teacher model to demonstrate correct behavior, and an accurate simulation environment to test in. In a company without a high-fidelity sandbox and reliable deterministic verifiers, building the execution layer is itself a serious engineering project. AutoSynthData does not remove that work. It changes what the work is for.

The teacher-model catch

There is a dependency in this design that deserves its own sentence. The whole pipeline runs on the gap between a weak model and a strong one. In the experiments, the teacher was a much larger model than the target. That works when you have a frontier system to borrow from. It works less well when you are the frontier, or when the task is so specialized that no stronger model exists to demonstrate it. In that case the pipeline has nothing to learn from, and the capability gap it depends on is simply a wall. Research on generating training data always runs into this eventually: the method scales as long as someone, somewhere, has already solved the problem.

Why this is the shape of enterprise AI

The interesting thing about this release is what it admits about where enterprise AI actually stands. The bottleneck is no longer getting a capable model. The bottleneck is getting a capable model to operate correctly inside a particular organization, with its particular tools and rules, where the failures are subtle and the state space is large.

For years the answer was manual annotation, which is slow, expensive and hard to scale. AutoSynthData is an argument that the failures themselves contain the signal, if you can extract, verify and diversify them without leaking the evaluation set. The pipeline is available for research on Hugging Face, and the team says reinforcement learning using the same methodology is next.

If it works broadly, the implication is that the scarce skill in enterprise AI shifts. It stops being about acquiring data and becomes about constructing environments good enough to generate trustworthy data inside. That is a harder problem, and a more durable one, because a company's environment is the thing no competitor can copy.

Related articles