🚀NEW LABGetting Started with Claude AgentsStart lab
Data · Agents · Training

1,000 Synthetic Computers

First page
1,000 Synthetic Computers
Paper summary

Microsoft Research builds 1,000 synthetic computers, each with realistic directory structures, documents, and artifacts, then runs long-horizon simulations on top of them. One agent plays the user and sets productivity goals; another executes the work. Each simulation runs over 8 hours of agent runtime and 2,000+ turns on average, roughly a month of human work compressed into one trace. Training on this experiential data drives significant improvements on both in-domain and out-of-domain productivity evaluations.

Ask this paper

Key points
01

Realistic synthetic environments: Each of the 1,000 computers ships with directory structures, documents, and artifacts that approximate a real user's working environment. The realism is what makes the trajectories useful as training data instead of as evaluation curiosities.

02

Two-agent simulation loop: A user agent sets productivity goals while a worker agent executes against them. The structure produces multi-turn, goal-directed traces that look like real productivity work, not the short scripted tasks that dominate existing benchmarks.

03

Designed to scale to billions of worlds: The framework is explicitly designed to scale to millions or billions of synthetic user worlds, which matches the scale at which frontier computer-use agents will need experiential data. The bottleneck on long-horizon training is data, and this is a credible recipe for producing it.

04

Why it matters: The bottleneck on computer-use agents has stopped being model capability and become realistic long-horizon training data. Synthetic-environment scaling is one of the few paths that does not depend on collecting massive amounts of real user telemetry, which makes it a practical default for teams building computer-use stacks.

Every Monday
Get next week’s papers.
Subscribe on Substack