Agent-Editing World Model: Rethinking World Modeling for LLM Agents

Shuang Sun, Guoxin Chen, Wayne Xin Zhao, Ji-Rong Wen and colleagues at Renmin University propose the Agent-Editing World Model (AEWM), which predicts how an agent's reasoning and actions affect task progress instead of predicting tool outputs.
Ask this paper
Target problem. Predicting high-entropy tool responses adds little when real feedback exists, while unsupported assumptions and stale plans stay in the history and distort later decisions (task-state contamination).
Action Judge. Classifies each decision as Critical, Exploratory or Noisy.
State Revision and EditAct. Noisy reasoning-action continuations are rewritten from the same observed history, and EditAct applies these edits during real execution so later decisions start from a corrected state.
Results. AEWM reaches 70.5% macro-F1 on the Action Judge benchmark, 10.6 points above the strongest frontier model; across six benchmarks and three backbones in search, terminal and SWE, EditAct adds 3.2 to 6.7 points over the strongest baseline.
Offline use. Rejection-sampling fine-tuning on verified EditAct trajectories beats Self-RFT by 2.2 to 2.6 points without needing AEWM online.
Abstract
Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from \emph{task-state contamination}, where unsupported assumptions and outdated plans persist in history and distort subsequent decisions. We propose the \textbf{Agent-Editing World Model (AEWM)}, which models how reasoning and actions shape future task progress rather than simulating tool responses. AEWM combines \textbf{Action Judge} to distinguish \textsc{Critical}, \textsc{Exploratory}, and \textsc{Noisy} decisions with \textbf{State Revision} to edit noisy reasoning--action continuations from the same observed history. \textbf{EditAct} integrates these capabilities with real execution, directly changing the state underlying subsequent decisions rather than merely providing critiques. We train AEWM across Search, Terminal, and Software Engineering through mid-training and supervised fine-tuning. AEWM achieves 70.5\% macro-F1 on our Action Judge benchmark, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2--6.7 points over the strongest baseline. Furthermore, rejection sampling fine-tuning on verified EditAct trajectories, termed \textbf{AEWM-RFT}, improves over Self-RFT by 2.2--2.6 points across three domains without online AEWM guidance.