Neal Desai Staff PM at Scale AI · ex-CTO · evals & agent tooling GH LI

Methodology · Synthetic data · Agent evals

Synthetic Artifacts and Evaluation Environments: Testing Agents on Work That Looks Real

Test agents on your real workflows without touching real data or real systems.

The quality loop Synthetic MCP servers Cost and reuse

What it does

Why you'd care

The artifact quality loop. A maker agent writes a file. Six checks run, by code and an independent model: valid, faithful, continuous, in voice, realistic, and novel. Realistic fails because the file reads like a template. Only that check is sent back, the maker repairs the file, and it is accepted with a receipt of every check.
The core of the method. A maker writes, independent checkers decide, and only what failed is repaired.

01Why enterprises need synthetic environments

An agent that files tickets, answers customers, or reconciles invoices touches several systems in one task. Before it runs on real work, someone has to know how it behaves on the messy, specific workflows of that business. Public benchmarks don't answer that question. They measure a generic distribution, they leak into training over time, and their answer keys are often wrong. Audits keep finding flawed tests in most of the SWE-bench Verified problems frontier models failed1, questionable answers in roughly a third of Humanity's Last Exam chemistry and biology items2, and errors in about 6.5% of MMLU3.

A company's own data can't fill the gap either. It is private, an agent under test could change it, and nobody has written down the right answers. What an enterprise needs is a synthetic copy of the kind of work its agents will do, with realistic artifacts, working tools, and known answers. The hard part is making all of that trustworthy.

02Start with the quality loop

Everything in this approach rests on one loop, run for every file, every world, and every test. A maker agent produces something. Code and a different model check it. If a check fails, only that check goes back for repair, a limited number of times. Then the result is accepted with a receipt of every check, or rejected with its reason. The maker never approves its own work, because models reliably favor what they wrote themselves4.

For an artifact, six checks run, cheapest first.

Files that pass one by one can still add up to something unrealistic, so the whole set is checked too. Its lengths and formats should match real examples, each kind of mess should appear at its measured rate, and the variety should stay close to real data. Variety needs particular care, because models converge on each other's writing so strongly that switching vendors buys little10. It has to come from a declared grid of personas, roles, file types, and mess levels.

Build note Keep every rejection reason. Counted up, they tell you exactly what to fix in the spec and the prompts, the same way reviewers' corrections sharpen a rubric5.

03One world under every artifact

Artifacts made one at a time drift apart. The fix is to build a world first, a structured record of who exists, what they own, what happened, and when, and render every artifact from it. The easiest worlds to make believable grow out of one person's working life, such as a sales lead, a support manager, or a founder. Their role, tools, and collaborators decide what exists. Answers to any question about the world are then queries on that record, so they are computed, never guessed. Recent work on synthetic companies has converged on the same design, sampling the organization first and rendering documents from it6, and computing answer keys from an entity graph7.

One event, the launch moving from May 1 to May 15, fans out to a meeting transcript, Slack, Notion, an email, and Linear. The checker reads each against the world graph. The email says May 12, which contradicts the graph, so only the email is repaired. Linear still says May 1, but that was planted on purpose and tagged, so it passes as drift. Result: five of five continuous, one drift tagged.
One event, every app. A real contradiction is caught and repaired; planted drift passes because it is tagged.

Continuity is what makes a world feel real. A decision in Tuesday's meeting should reach the Slack channel within the hour, the project page by Wednesday, and the client email after that. The check works in five layers: within a file, against the world, across apps, over time, and per person, so that nobody quotes a decision before it was made or signs off on something above their role. Structured fields are parsed by code. Free text is broken into small, checkable claims by a model other than the author, as FActScore does8, and each claim is looked up in the world9. The target is zero contradictions that nobody planted.

Build note Build the continuity checker early, and test it on a small hand-made world with errors you planted yourself. You need to know its accuracy before it judges anything else.

04Copy the mess on purpose

Real company data is messy in consistent ways, and agents fail on exactly that mess. A world that is too tidy flatters every agent tested in it.

stale versionsnear-duplicatesconflicting numbersmissing attachmentsmisfiled documentsbroken formatstypos and shorthandaccess gapsdates that slip

The mess is measured in real, license-cleared sources, reproduced at the measured rates, and tagged so nothing downstream mistakes it for an error. The tag also makes it useful. A planted discrepancy is a question with a known answer, such as which launch date is current.

05Synthetic MCP servers

Artifacts test what an agent reads. Tools test what it does. Agents increasingly reach their tools through MCP servers, so the environment should give them MCP servers too. Each real server the enterprise uses (Slack, Gmail, Drive, Salesforce, Jira) gets a synthetic stand-in that the agent calls in exactly the same way.

Agentunder test Synthetic MCP serverSlack · Gmail · Drive · CRM real MCP serverswap in later, agent unchanged State storeloaded from the world Grader APIhidden from the agent same tools, same arguments and errors reads state and the call log
The agent sees the real tool surface. Behind it sits state loaded from the world, and a grader the agent cannot reach.

What the stand-in has to mimic is more than the happy path.

MimicWhy it matters
Tool names and schemasThe agent's tool calls should transfer unchanged to the real server
Pagination and searchAgents that read only the first page miss what matters
Permissions and scopesReal users can't see everything, and neither should the agent
Errors and rate limitsRecovery from failure is part of the job
PreconditionsSome actions only work after others, like sharing a file before mentioning it
Side effectsSending a message should change what others see next

Three design choices matter most. First, state lives in code. Models write the message bodies and documents when the world is built, but a real store holds the state, because models simulating state break down as soon as several fields must change together11. The strongest environments are all built this way1213. Second, realistic edges, because tools that only succeed inflate every score; implicit preconditions are a known place where agents fail15. Third, a hidden grader side that reads the true state and the log of every call, so a task can be scored on its outcome. Recorded responses from the real server can fill gaps for read-heavy tools14, and each task starts from a known, seeded state and ends with a verification script16.

Build note Keep the synthetic server's tool surface contract-tested against the real server's published schemas. When they match, the same agent and the same tests run against staging or production later with no changes.

06Write the test from the answer

It is easier to build a hard problem while holding the answer than to solve it from scratch without one. A maker that builds a test already knows its solution, because it planted the fact, broke the code, or wrote the record the question depends on. The model under test starts blind. AutoBencher gives its generator privileged information for the same reason17, and SWE-smith breaks working code so the fix is known in advance18.

Pick theanswer Buildbackwards Replay thegolden path Keyedsolver Blindpanel Rubricfrom the paths from the worldmaker's pathreaches the gold?item sound?how hard?expert reviews
The maker's path is the golden trajectory; independent solvers add silver ones; the rubric is written from both.

The same approach covers questions whose answers are spread across apps, act-or-ask decisions like those in Escalation Bench, goals the agent must reach through the synthetic servers (graded on the end state, with a check that nothing else changed), and work products such as a memo or a spreadsheet.

07What it costs, and why reuse matters

None of this is cheap. The loop only works with frontier models in the seats that matter. They write artifacts that pass for real, extract claims reliably, judge realism, and solve the hardest items. Every accepted artifact costs its generation, several checks by more than one model, any repairs, and a share of every artifact that was rejected. Every accepted test item adds a keyed solver and a blind panel from several labs, run multiple times. Design takes expert time too, in writing specs, grading pilots, and calibrating judges. The market prices this accordingly. Industry reporting puts single agent tasks at hundreds to a couple of thousand dollars each, a replica of one website around twenty thousand, and a complex product simulation up to three hundred thousand19.

Three levers keep the cost sane. Run checks cheapest first, so code rejects what it can before any model is called. Repair only what failed, with a cap on rounds. And above all, make as much as possible reusable.

Build note Track cost per accepted artifact and per accepted test, including rejections. It shows where a cheaper model or an earlier code check can take a seat without hurting quality.

08What to build, in order

Artifact forge

Writes realistic files in native formats, run through the quality loop.

Continuity checker

Audits any set of artifacts against the world, including sets made elsewhere.

World builder

Builds the record of people, events, and time from a description of one person's work.

Profiler

Measures how real documents are shaped and how they are messy.

Synthetic MCP servers

Mimic real tool servers over state loaded from the world.

Test builder and graders

Build tests from their answers, and grade on outcomes first.

  1. The artifact forge and its quality loop. Useful on its own, and everything else depends on it.
  2. The continuity checker, proven on a small world with planted errors.
  3. A world builder that starts from one person and rebuilds identically from a seed.
  4. Measured mess, profiled from real sources.
  5. Synthetic MCP servers for the enterprise's most-used tools, contract-tested against the real ones.
  6. Tests and graders, answers first.

People stay in the loop where quality is decided. They approve each world and each eval's definition, write the rubric by grading about thirty pilot items, calibrate every automated judge against their own labels, review the hardest items and the disagreements, and audit a sample of everything before release.

One idea runs through all of it. Whatever makes something never gets the last word on it. The world checks the files, the files check each other, the real servers' contracts check the synthetic ones, independent models check the tests, and people check the checkers. That chain of receipts is what an enterprise needs before it lets an agent near real work.


Companion pieces: Escalation Bench, an agentic benchmark built with an earlier version of this loop; Loop Engineering, on measuring agent skills; and Relay, an agent that works through tools much like the ones simulated here.

Sources

  1. OpenAI, Why we no longer evaluate SWE-bench Verified (2026).
  2. Phan et al., Humanity's Last Exam (2025); FutureHouse, audit of HLE chemistry and biology answers.
  3. Gema et al., Are We Done with MMLU? (2024).
  4. Xu et al., When LLMs Benchmark Themselves (2025).
  5. Shankar et al., Who Validates the Validators? (2024).
  6. Hamilton et al., CorporateBench (2026).
  7. Gruenbaum et al., The Era by Eon Benchmark (2026).
  8. Min et al., FActScore (2023).
  9. Wei et al., Long-form factuality in large language models (2024).
  10. Jiang et al., Artificial Hivemind (2025).
  11. Liu et al., EnvSimBench (2026).
  12. Yao et al., τ-bench (2024).
  13. Trivedi et al., AppWorld (2024).
  14. Guo et al., StableToolBench (2024).
  15. Lu et al., ToolSandbox (2024).
  16. Wu et al., MCPMark (2025).
  17. Li et al., AutoBencher (2024).
  18. Yang et al., SWE-smith (2025).
  19. Epoch AI, An FAQ on RL environments (2026).