Methodology · Synthetic data · Agent evals
Synthetic Artifacts and Evaluation Environments: Testing Agents on Work That Looks Real
Test agents on your real workflows without touching real data or real systems.
What it does
- Builds company data that passes for real. Slack threads, email, documents, and CRM records that agree with each other.
- Serves it through stand-in tool servers. Synthetic MCP servers that look and behave like Slack, Gmail, Drive, or Salesforce to the agent.
- Writes tests with known answers. Every answer is checked before anyone relies on it.
Why you'd care
- See before you ship. Know how an agent handles your workflows before it touches production.
- Nothing at risk. No private data leaves, and no real system can be changed.
- Scores you can trust. Every file and test is checked by something other than what made it.
- Built once, reused often. Doing this well is expensive, but worlds, servers, and personas carry over to every new eval.
01Why enterprises need synthetic environments
An agent that files tickets, answers customers, or reconciles invoices touches several systems in one task. Before it runs on real work, someone has to know how it behaves on the messy, specific workflows of that business. Public benchmarks don't answer that question. They measure a generic distribution, they leak into training over time, and their answer keys are often wrong. Audits keep finding flawed tests in most of the SWE-bench Verified problems frontier models failed1, questionable answers in roughly a third of Humanity's Last Exam chemistry and biology items2, and errors in about 6.5% of MMLU3.
A company's own data can't fill the gap either. It is private, an agent under test could change it, and nobody has written down the right answers. What an enterprise needs is a synthetic copy of the kind of work its agents will do, with realistic artifacts, working tools, and known answers. The hard part is making all of that trustworthy.
02Start with the quality loop
Everything in this approach rests on one loop, run for every file, every world, and every test. A maker agent produces something. Code and a different model check it. If a check fails, only that check goes back for repair, a limited number of times. Then the result is accepted with a receipt of every check, or rejected with its reason. The maker never approves its own work, because models reliably favor what they wrote themselves4.
For an artifact, six checks run, cheapest first.
- Valid. It opens and parses like the real format. A Slack export loads, and a spreadsheet's formulas compute.
- Faithful. Every fact traces to the world it came from, or to mess that was planted on purpose.
- Continuous. It agrees with every other artifact about the same facts.
- In voice. It sounds like its author, in the register of its medium: terse in Slack, formal in a contract.
- Realistic. A judge shown it beside real examples cannot reliably tell which is which.
- Novel. It is more than a near-copy of something already made, unless the duplicate is deliberate.
Files that pass one by one can still add up to something unrealistic, so the whole set is checked too. Its lengths and formats should match real examples, each kind of mess should appear at its measured rate, and the variety should stay close to real data. Variety needs particular care, because models converge on each other's writing so strongly that switching vendors buys little10. It has to come from a declared grid of personas, roles, file types, and mess levels.
Build note Keep every rejection reason. Counted up, they tell you exactly what to fix in the spec and the prompts, the same way reviewers' corrections sharpen a rubric5.
03One world under every artifact
Artifacts made one at a time drift apart. The fix is to build a world first, a structured record of who exists, what they own, what happened, and when, and render every artifact from it. The easiest worlds to make believable grow out of one person's working life, such as a sales lead, a support manager, or a founder. Their role, tools, and collaborators decide what exists. Answers to any question about the world are then queries on that record, so they are computed, never guessed. Recent work on synthetic companies has converged on the same design, sampling the organization first and rendering documents from it6, and computing answer keys from an entity graph7.
Continuity is what makes a world feel real. A decision in Tuesday's meeting should reach the Slack channel within the hour, the project page by Wednesday, and the client email after that. The check works in five layers: within a file, against the world, across apps, over time, and per person, so that nobody quotes a decision before it was made or signs off on something above their role. Structured fields are parsed by code. Free text is broken into small, checkable claims by a model other than the author, as FActScore does8, and each claim is looked up in the world9. The target is zero contradictions that nobody planted.
Build note Build the continuity checker early, and test it on a small hand-made world with errors you planted yourself. You need to know its accuracy before it judges anything else.
04Copy the mess on purpose
Real company data is messy in consistent ways, and agents fail on exactly that mess. A world that is too tidy flatters every agent tested in it.
The mess is measured in real, license-cleared sources, reproduced at the measured rates, and tagged so nothing downstream mistakes it for an error. The tag also makes it useful. A planted discrepancy is a question with a known answer, such as which launch date is current.
05Synthetic MCP servers
Artifacts test what an agent reads. Tools test what it does. Agents increasingly reach their tools through MCP servers, so the environment should give them MCP servers too. Each real server the enterprise uses (Slack, Gmail, Drive, Salesforce, Jira) gets a synthetic stand-in that the agent calls in exactly the same way.
What the stand-in has to mimic is more than the happy path.
| Mimic | Why it matters |
|---|---|
| Tool names and schemas | The agent's tool calls should transfer unchanged to the real server |
| Pagination and search | Agents that read only the first page miss what matters |
| Permissions and scopes | Real users can't see everything, and neither should the agent |
| Errors and rate limits | Recovery from failure is part of the job |
| Preconditions | Some actions only work after others, like sharing a file before mentioning it |
| Side effects | Sending a message should change what others see next |
Three design choices matter most. First, state lives in code. Models write the message bodies and documents when the world is built, but a real store holds the state, because models simulating state break down as soon as several fields must change together11. The strongest environments are all built this way1213. Second, realistic edges, because tools that only succeed inflate every score; implicit preconditions are a known place where agents fail15. Third, a hidden grader side that reads the true state and the log of every call, so a task can be scored on its outcome. Recorded responses from the real server can fill gaps for read-heavy tools14, and each task starts from a known, seeded state and ends with a verification script16.
Build note Keep the synthetic server's tool surface contract-tested against the real server's published schemas. When they match, the same agent and the same tests run against staging or production later with no changes.
06Write the test from the answer
It is easier to build a hard problem while holding the answer than to solve it from scratch without one. A maker that builds a test already knows its solution, because it planted the fact, broke the code, or wrote the record the question depends on. The model under test starts blind. AutoBencher gives its generator privileged information for the same reason17, and SWE-smith breaks working code so the fix is known in advance18.
- The gold answer is computed from the world and re-derived by code. Knowing an answer is not the same as it being right, and items built to be hard fail audits more often.
- A keyed solver, given the answer key, tests the item. If it still can't get there, the item is ambiguous or broken.
- A blind panel of models from several labs tests the models. Its solve rate is the difficulty, and its successes become extra valid paths, so the grader doesn't punish a route the maker didn't take.
- The rubric comes from those paths. Steps every valid path shares become checkpoints, harmful moves blind solvers made become forbidden actions, and an expert reviews the result.
The same approach covers questions whose answers are spread across apps, act-or-ask decisions like those in Escalation Bench, goals the agent must reach through the synthetic servers (graded on the end state, with a check that nothing else changed), and work products such as a memo or a spreadsheet.
07What it costs, and why reuse matters
None of this is cheap. The loop only works with frontier models in the seats that matter. They write artifacts that pass for real, extract claims reliably, judge realism, and solve the hardest items. Every accepted artifact costs its generation, several checks by more than one model, any repairs, and a share of every artifact that was rejected. Every accepted test item adds a keyed solver and a blind panel from several labs, run multiple times. Design takes expert time too, in writing specs, grading pilots, and calibrating judges. The market prices this accordingly. Industry reporting puts single agent tasks at hundreds to a couple of thousand dollars each, a replica of one website around twenty thousand, and a complex product simulation up to three hundred thousand19.
Three levers keep the cost sane. Run checks cheapest first, so code rejects what it can before any model is called. Repair only what failed, with a cap on rounds. And above all, make as much as possible reusable.
- Worlds. One verified world supports many evals: questions, tool tasks, and decisions all built on the same people and history. Held-out tests get fresh worlds so answers can't leak.
- Synthetic servers. A Slack or CRM stand-in is built and contract-tested once, then loaded with any world.
- Persona libraries. Well-made personas, with voice, habits, and role, carry over to new worlds in the same industry.
- Structure and mess profiles. Once measured for a file type, they apply to every world that uses it.
- Verified artifacts. Accepted files become exemplars for the realism judge and templates for the next world.
- Variants. Changing one fact in a world and re-rendering only the files it touches produces a minimal-pair variant, and only those files need re-checking.
Build note Track cost per accepted artifact and per accepted test, including rejections. It shows where a cheaper model or an earlier code check can take a seat without hurting quality.
08What to build, in order
Artifact forge
Writes realistic files in native formats, run through the quality loop.
Continuity checker
Audits any set of artifacts against the world, including sets made elsewhere.
World builder
Builds the record of people, events, and time from a description of one person's work.
Profiler
Measures how real documents are shaped and how they are messy.
Synthetic MCP servers
Mimic real tool servers over state loaded from the world.
Test builder and graders
Build tests from their answers, and grade on outcomes first.
- The artifact forge and its quality loop. Useful on its own, and everything else depends on it.
- The continuity checker, proven on a small world with planted errors.
- A world builder that starts from one person and rebuilds identically from a seed.
- Measured mess, profiled from real sources.
- Synthetic MCP servers for the enterprise's most-used tools, contract-tested against the real ones.
- Tests and graders, answers first.
People stay in the loop where quality is decided. They approve each world and each eval's definition, write the rubric by grading about thirty pilot items, calibrate every automated judge against their own labels, review the hardest items and the disagreements, and audit a sample of everything before release.
One idea runs through all of it. Whatever makes something never gets the last word on it. The world checks the files, the files check each other, the real servers' contracts check the synthetic ones, independent models check the tests, and people check the checkers. That chain of receipts is what an enterprise needs before it lets an agent near real work.
Companion pieces: Escalation Bench, an agentic benchmark built with an earlier version of this loop; Loop Engineering, on measuring agent skills; and Relay, an agent that works through tools much like the ones simulated here.
Sources
- OpenAI, Why we no longer evaluate SWE-bench Verified (2026).
- Phan et al., Humanity's Last Exam (2025); FutureHouse, audit of HLE chemistry and biology answers.
- Gema et al., Are We Done with MMLU? (2024).
- Xu et al., When LLMs Benchmark Themselves (2025).
- Shankar et al., Who Validates the Validators? (2024).
- Hamilton et al., CorporateBench (2026).
- Gruenbaum et al., The Era by Eon Benchmark (2026).
- Min et al., FActScore (2023).
- Wei et al., Long-form factuality in large language models (2024).
- Jiang et al., Artificial Hivemind (2025).
- Liu et al., EnvSimBench (2026).
- Yao et al., τ-bench (2024).
- Trivedi et al., AppWorld (2024).
- Guo et al., StableToolBench (2024).
- Lu et al., ToolSandbox (2024).
- Wu et al., MCPMark (2025).
- Li et al., AutoBencher (2024).
- Yang et al., SWE-smith (2025).
- Epoch AI, An FAQ on RL environments (2026).