Breaking the Environment Wall: Evolving LLM Agent Environments for Recursive Self-Improvement

Yukai Wu, Xuanhe Zhou, Fan Wu and colleagues at Shanghai Jiao Tong University, Theseus Labs and Tencent Hunyuan present Env-Rethink, a system with a 27B post-trained model that reorganizes messy work environments for agents and then evolves them into harder ones.
Ask this paper
Problem. Real environments scatter related information across files, mix evidence with misleading or outdated versions, and change over time. Averaged over nine agent configurations, pass rates fall from 83.9% in clean environments to 57.6% in noisy ones, a drop of 26.3 points.
Organization. Collection Maps group related files and Event Logs record cross-file relationships, so an agent gets the context it would otherwise have to reconstruct.
Learned preparation. A 27B model trained on trajectories learns which files to read and how to verify cross-file claims, then hands a separate downstream agent the selected files and an evidence report. Across nine downstream models on 30 tasks, mean rubric pass rate rises from 57.6% with full environments and 59.4% with a Qwen3.8-27B preparer to 72.7%, with gains for every model.
Environment evolution. Virtual event histories change environment state and evidence relationships under a fixed task request. On Terminal-Bench 2.1, the evolved versions lower success for at least three models on 32 of 55 tasks, which supplies harder tasks for further agent training.
Abstract
Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, relevant evidence in the environment is often mixed with misleading information and conflicting versions. Third, environments evolve over time, introducing new noise and more challenging tasks. These challenges can substantially degrade performance for state-of-the-art AI agents (e.g., from 83.9% to 57.6%). To address these challenges, we propose Env-Rethink (a system with 27B post-trained model) that supports three main capabilities: (1) It adaptively builds Collection Maps (for organizing related files) and Event Logs (for contextualizing cross-data relationships) to supplement necessary context; (2) It further leverages the post-trained model (through offline trajectory learning) to identify underlying noise issues in the environment; (3) It ultimately evolves environments through virtual event histories that alter environmental states and evidence relationships, producing more tricky ones for further agent improvement. Experiments show that Env-Rethink can effectively improve downstream task performance (with over 15.1% rubric pass rate improvement across nine models on 30 tasks).