Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Jie Wu and colleagues on the Qwen team at Alibaba with Tsinghua turn the pile of existing terminal-agent trajectories into executable environments, on the observation that a trajectory's tool-execution history already exposes the environment it ran in.
Ask this paper
Trajectories are frozen, environments are re-queryable: an environment can be turned into many verifiable tasks and returns execution feedback, while a trajectory is a single demonstration. Post-training needs the former and the community has accumulated the latter.
Reconstruct rather than synthesize: Terminal-Universe replays the file operations recorded in a trajectory to restore each file to its pre-modification state, yielding a partial workspace, then a completion agent supplies the missing files and dependencies.
Scaled on two axes: for breadth it mines directional dependency relations between related environments and synthesizes cross-workspace queries spanning multiple codebases, and for depth it extends a single-turn query into a multi-round session with a user agent supplying iterative feedback.
37.3k task-sufficient environments produced from public terminal agent trajectories, which is the scale number that makes the approach interesting rather than merely clever.
Real downstream gains: supervised fine-tuning of Qwen3.5-27B on the corpus improves Terminal-Bench 2.1 single-round by 11.9 points and EvoCode-Bench v2 MT@4 multi-round by 13.8 points.
Abstract
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure and contents of the environments in which they ran, making it possible to reconstruct those environments from the trajectories themselves. Thus, we introduce Terminal-Universe, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions. Specifically, Terminal-Universe replays the file operations recorded in a trajectory to restore each file before the agent modified it, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On this recovered workspace, we both reconstruct the original intent task and synthesize entirely new ones. Besides, we also scale the tasks along two complementary axes: breadth and depth. For breadth, we mine directional dependency relations between related environments and synthesize cross-workspace queries spanning multiple codebases, as developers routinely do in real-world development. For depth, we extend the initial single-turn query into a multi-round session that captures iterative user feedback and requirement refinement via a user agent. Applied to public terminal agent trajectories, Terminal-Universe produces 37.3k task-sufficient environments. Supervised fine-tuning of Qwen3.5-27B on this corpus improves single-round performance on Terminal-Bench 2.1 by 11.9 points and multi-round performance on EvoCode-Bench v2 MT@4 by 13.8 points.