AutoSynthData: Generating Training Data for Enterprise Agents
AutoSynthData: Generating Training Data for Enterprise Agents Published October 2, 2026 Enterprise agents must operate effectively within the systems, policies, workflows, and d...
By AI Engineering Team
AutoSynthData: Generating Training Data for Enterprise Agents
Published October 2, 2026
Enterprise agents must operate effectively within the systems, policies, workflows, and data environments of the organizations that use them. A model can be broadly capable while still struggling with a specific workflow, tool combination, or operational constraint.
The challenge is converting those weaknesses into useful training data. One failure reveals a capability gap, but training requires many varied tasks that exercise the same capability. Each task must be executable in the target environment, resemble a realistic request, and include a dependable way to determine whether the agent succeeded.
ServiceNow CoreAI developed AutoSynthData to address this problem. The system uses failures from a target model and successful behavior from a stronger teacher model to identify what the target should learn next. It then generates and validates new tasks focused on those capabilities. As the target improves, the curriculum shifts toward the remaining weaknesses.
The approach is illustrated with EnterpriseOps Gym (Malay et al., 2026) and the released dataset.
What Makes a Useful Agentic Task?
An agentic environment defines the world in which an agent operates. It includes the state the agent can observe and modify, the tools and APIs it can invoke, and the state transitions caused by its actions.
A task is instantiated within that environment as:
task = (system specification, user prompt, verifier)
System Specification
The system specification defines the constraints under which the agent operates. It can include system instructions, environment policies, and task-specific initialization, such as a seeded database state or a collection of knowledge articles.
The specification must be compatible with the environment's tools, state, and supported actions. Its instructions should be clear and should not introduce arbitrary constraints solely to make a task more difficult.
Agent-Facing Task
The user prompt describes what the agent must accomplish and any user-level constraints. A generated task should meet three requirements:
- Feasibility: At least one trajectory in the current environment must satisfy the prompt while following the system specification. This excludes tasks that require unavailable tools, inaccessible knowledge, impossible state transitions, or actions prohibited by policy.
- Realism: The request should resemble something a user could plausibly ask in the target environment. The set of executable behaviors is generally much larger than the set of realistic workflows.
- Difficulty: The task should reveal a weakness in the current agent. Tasks that the model already solves reliably provide limited training value. The useful region consists of tasks that are feasible and realistic but not yet consistently solved.
Verifier
The verifier determines whether the resulting trajectory completes the task successfully. It should provide:
- Consistency: It should agree with the user prompt, system specification, and task-specific environment state.
- Soundness: It should reject trajectories that fail to satisfy the task or violate relevant constraints.
- Completeness: It should accept valid solutions rather than requiring one specific reference trajectory.
These properties directly affect training. A lax verifier can reward incorrect behavior, while an overly restrictive verifier can reject valid solutions.
AutoSynthData Overview
Given an environment and a target model, AutoSynthData generates training tasks containing a system specification, user prompt, and verifier. The tasks are grounded in the environment and selected to provide useful training signal for the current model.
The process begins by evaluating the target model with diagnostic tasks and identifying patterns in its failures. A stronger teacher helps determine which tasks are solvable and what successful behavior looks like. AutoSynthData converts the resulting capability gaps into new executable tasks, validates each task in the environment, and uses accepted samples for post-training. Evaluation of the updated model then reveals which gaps remain and informs the next generation cycle.
From Model Failures to a Curriculum
AutoSynthData uses evaluation runs in the target environment to identify what the model needs to learn next. In the EnterpriseOps Gym experiment, both the target model and a stronger teacher run on evaluation tasks. The results are examined to identify:
- the capability being tested;
- the tools and workflow structure involved;
- where the target model fails and how the teacher succeeds;
- the properties a correct final state must satisfy;
- the dimensions that can vary while preserving the capability being tested.
These findings are distilled into sanitized capability specification cards. The evaluation tasks indicate what the model should learn, but the generator does not receive their original prompts, entities, trajectories, or verifier details. Instead, it receives the cards and uses them to create new tasks with different prompts, states, and solution paths.
Generating and Scaling Tasks
A capability gap identifies what to teach, but training requires many varied tasks that exercise that capability. AutoSynthData uses each specification card to generate those tasks.
For example, if the target model struggles with a particular workflow, the generator can create tasks that exercise the same workflow while varying:
- entities;
- initial environment state;
- workflow composition;
- tool combinations;
- wording;
- difficulty.
The stronger teacher demonstrates a successful trajectory for each task. For supervised fine-tuning (SFT), these demonstrations teach the target model to apply the capability in new situations.
AutoSynthData builds the dataset in two phases: generating and validating core samples, then expanding accepted samples into novel variants.
Target
The target phase creates the core training set from the capability specifications. Workers generate independent tasks in parallel and begin a new target when they finish. Each candidate passes through validation, execution, solver evaluation, and repair before acceptance. The result is a set of vetted examples focused on the target model's learning needs.
Multiply
The multiply phase expands the dataset by creating novel variants of accepted target samples. Each variant has its own user request, environment state, entity configuration, reference trajectory, and verifier. It must pass the same validation and execution checks.
A multiplied sample cannot seed another multiplied sample. This keeps expansion anchored to the vetted target set and limits drift across generations.
Implementation Details
AutoSynthData separates generation control from environment-specific execution. A shared controller coordinates generation, quality control, coverage, and dataset construction. An adapter manages environment execution, task and state management, reference replay, deterministic verification, solver execution, and task profiling.
Parallel target generation and multiplication provide a path to training-scale datasets. Their usefulness depends on checking every candidate: the task must be executable, the solution must work, and the verifier must distinguish success from failure.
High-Quality Synthetic Data Requires More Than Generation
A plausible request alone is not sufficient for useful training data. A task may be impossible in the target environment, its reference solution may fail during execution, or its verifier may reward the wrong final state. AutoSynthData checks these properties before accepting a task.
Quality is reviewed at two levels. Individual candidates must pass verification, and batches must provide useful coverage and diversity.
Sample-Level Verification and Repair
Every candidate passes through a quality-control loop before entering the training dataset. Solver evaluation first measures difficulty. In the described configuration, tasks are favored when the target model solves them on no more than one of three trials and the stronger solver solves them on at least two of three trials.
Candidates also undergo positive verification, negative verification, and bounded repair.
Positive Verification
The positive gate asks whether the intended solution completes the generated task.
The pipeline executes the reference trajectory in the target environment and compares the resulting state with the candidate's verifier. This exposes mismatches among the prompt, initial state, solution, and success criteria.
Negative Verification
The negative gate asks whether relevant incorrect outcomes fail.
For example, parts of the expected outcome can be mutated to confirm that those states no longer pass verification. This identifies weak verifiers that award success without requiring the intended behavior.
Critique and Repair
Failed candidates are sent to a critic before they are discarded. The critic examines the sample and its failure, looking for inconsistent state, impossible workflows, incorrect task construction, faulty reference trajectories, weak verifier logic, or a mismatch with the intended capability.
The findings guide repairs with a fixed limit on retries:
candidate
↓
failure
↓
critique / diagnosis
↓
targeted repair
↓
run the gates again
↓
accept or retry
A repaired task must pass the relevant checks again. The diagnosis guides repairs to the existing candidate rather than requiring generation to start over.
Passing these checks makes a sample eligible for training, but individually valid samples can still produce a repetitive or unbalanced dataset. AutoSynthData therefore reviews batches as well.
Batch-Level Review
A batch may overrepresent a few easy task families, omit a capability, or devote too much generation effort to a low-yield pattern.
A meta-review examines accepted samples, rejected samples, and generation behavior across each batch. It asks:
- Which task families are overrepresented, and which capability dimensions are missing?
- Are the same types of examples appearing repeatedly?
- Do particular targets repeatedly fail generation?
- Are systematic problems appearing in critiques?
- What guidance should change for the next batch?
The controller tracks coverage in the accepted dataset, reduces generation in overrepresented regions, and directs more work toward gaps. When a region repeatedly produces poor candidates, critiques and meta-review guide changes to the generation strategy.
These adjustments balance learning signal, task quality, coverage, diversity, redundancy, generation budget, and dataset-size requirements. Sample-level checks improve individual candidates, while batch-level review guides future generation.
Moving the Training Frontier
The most useful training distribution changes as the model improves. AutoSynthData treats synthetic data generation as a search for tasks near the target model's capability boundary: tasks difficult enough to expose weaknesses, but solvable enough for the teacher to provide reliable demonstrations.
After post-training, the updated model is evaluated in the same environment. Tasks it now solves reliably are less useful for the next training round, while persistent failures identify capabilities that still need attention. These results guide the next generation cycle.
The experiments described here focus on SFT, but the same mechanism could support reinforcement learning (RL). The system could generate tasks that challenge the current policy, provide reliable learning signals, train the policy, and then move the generation target as the policy changes. The planned direction is to test this moving, difficulty-calibrated frontier beyond SFT.
EnterpriseOps Gym Experiments
EnterpriseOps Gym was used to test whether the approach improves a model on tasks in a stateful enterprise environment. Training tasks were generated in the Gym's Hybrid and ITSM environments, the target model was fine-tuned on accepted samples, and the resulting checkpoints were evaluated.
Hybrid
The pipeline was tested on the Hybrid domain of EnterpriseOps Gym with Gemma-4-26B-A4B-it as the target model and Qwen3.8-27B as the teacher.
AutoSynthData generated 2,000 synthetic training samples in about 18 hours. Gemma was fine-tuned on this dataset, and the resulting checkpoints were evaluated on the benchmark. The best checkpoint was epoch 5.
Hybrid Results
The synthetic SFT checkpoint improved mean Pass@1 by 7.2 percentage points, a 35% relative improvement. Verifier success increased from 63.01% to 68.55%. The checkpoint closed 59% of the original Pass@1 gap between Gemma and the reference model.
The training tasks were newly generated from capability specifications. The generator did not receive the original evaluation tasks. The result demonstrates improvement in the EnterpriseOps Gym Hybrid environment used for the experiment.
ITSM
AutoSynthData was also applied to the ITSM domain of EnterpriseOps Gym, using Gemma-4-26B-A4B-it as the target model and DeepSeek-V4.1-Flash as the teacher.
The system generated 1,994 synthetic training samples in 66 hours. Generation took longer than in the subsequent Hybrid run, primarily because the ITSM run used a larger teacher model and took place before pipeline optimizations improved throughput.
In ITSM, synthetic SFT increased mean Pass@1 from 18.77% to 27.18%, indicating improvement in a second domain.
Closing the Loop
The tasks most useful for training depend on both the environment and the model operating within it. AutoSynthData uses model failures to select what to generate, validates new tasks against the environment, and makes accepted tasks available for post-training.
The EnterpriseOps Gym results demonstrate this process in a controlled setting. As the model changes, the same feedback loop can focus generation on the capability gaps that remain.