StandardsAboutContact
The Weights
ServiceNow's AutoSynthData: Wait for Replication

ServiceNow's AutoSynthData: Wait for Replication

ServiceNow CoreAI reports that about 2,000 synthetic, verifier-checked tasks lifted a Gemma-4-26B-A4B-it agent by 7.2 Pass@1 points on a hybrid enterprise benchmark. The method is sensible, but the evidence is one run reported by the organization that also maintains the benchmark.

Wait for independent replication before committing budget. ServiceNow CoreAI reports its AutoSynthData loop raised Gemma-4-26B-A4B-it by 7.2 Pass@1 points, a 35% relative gain, on hybrid enterprise tasks. The evidence is a vendor blog scored on a benchmark from the same organization, with no reported seeds or variance.

The Weights Desk · 4 min read

The verdict on ServiceNow CoreAI's AutoSynthData is Wait. In a blog dated October 2, 2026, ServiceNow CoreAI reports that a gap-targeted synthetic-data loop lifted Gemma-4-26B-A4B-it by 7.2 Pass@1 points on a hybrid enterprise task set, a 35% relative gain. The method is plausible, but the figure comes from a single vendor report on a benchmark from the same organization, and the blog reports no variance. We read the blog, the dataset card and the benchmark paper; we did not run the pipeline.

What the pipeline actually does

AutoSynthData treats training data as a response to measured failure rather than bulk generation. Per the blog, it runs the target model and a stronger teacher on diagnostic tasks, summarizes failure patterns as capability specification cards, generates new tasks to cover them, validates the results, fine-tunes on accepted samples and re-evaluates. Each task pairs a system specification and user prompt with a verifier, so correctness is checked by executable logic rather than by another model's opinion.

The reported numbers, and what they omit

On the hybrid domain, the blog reports 2,000 synthetic samples generated in about 18 hours, Pass@1 up 7.2 points, verifier success up from 63.01% to 68.55%, and 59% of the Pass@1 gap to a reference model closed, using Qwen3.8-27B as teacher. The blog states no seeds, variance or confidence intervals. We also found no plain-distillation baseline without gap targeting, so the benefit of the targeting step itself is not isolated.

ITSM shows the gain and the cost

In the IT service management run, the blog reports Pass@1 rising from 18.77% to 27.18% with 1,994 samples and DeepSeek-V4.1-Flash as teacher. Generation took 66 hours, against roughly 18 for the hybrid run. That is a gain of about 8.4 points, with the absolute score still under 30%. The blog attributes part of the time to the teacher and notes later optimizations, so treat the figure as a first-run cost, not a steady-state price.

The benchmark is rigorous but low-ceiling and self-owned

EnterpriseOps-Gym scores final environment state with executable SQL verifiers across eight domains and 512 tools, which is the right design. Its origins disagree on details: the paper cites 1,150 tasks and a top score of 37.4% (Claude Opus 4.5), while the dataset card lists 1,115 tasks in its main benchmark and a best score of 34.1%. Either way, ceilings are low. ServiceNow-AI maintains both the benchmark and the reported gains, so independent replication matters.

Who should use it, and who should not

Teams that already have executable verifiers for their workflows, such as database-state checks on ticketing or HR systems, are the natural users, because the method depends on them. Teams without verifiers, or those expecting production reliability from scores of roughly 27% or 69%, should not treat this as a fix. A pilot should build verifiers, measure a plain-distillation baseline, run several seeds and hold out tasks the generator never saw.

The bottom line

Wait. The deciding number is 7.2 Pass@1 points on a single reported hybrid run, with no stated variance and no outside replication. Targeting measured gaps and filtering with positive and negative checks is sound enough to pilot on your own tasks. It is not enough evidence to commit a training budget or promise reliability gains to stakeholders. The call changes if an outside team publishes multi-seed results.

What does AutoSynthData do?
It builds fine-tuning data for enterprise agents by finding where a target model fails relative to a stronger teacher, summarizing those gaps, generating new verifier-backed tasks, filtering them through quality checks, and retraining. ServiceNow CoreAI describes this in its Hugging Face blog.
How large is the reported improvement?
On the hybrid domain, the blog reports Pass@1 up 7.2 points (35% relative), verifier success from 63.01% to 68.55%, and 59% of the gap to a reference model closed. On ITSM, Pass@1 moved from 18.77% to 27.18%. The blog does not report seeds, variance or confidence intervals.
Should an enterprise team adopt it now?
Not as a drop-in. The results come from one vendor blog and a benchmark the same organization maintains. A small pilot on your own workflows, with held-out tasks, a plain-distillation baseline and several seeds, is the sensible first step.
  1. AutoSynthData: Generating Training Data for Enterprise Agents — Hugging Face (ServiceNow-AI)
  2. ServiceNow-AI/EnterpriseOps-Gym dataset card — Hugging Face (ServiceNow-AI)
  3. EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings — arXiv