🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 4, 2026
Agents

When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

First page
When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents
The curator’s take

Janvijay Singh (UIUC, Microsoft Research intern), Vaishnavi Shrivastava, Dilek Hakkani-Tür, Ece Kamar and Asli Celikyilmaz at Microsoft Research AI Frontiers introduce AGNI, a pipeline that breaks one environmental assumption behind a successful terminal-agent trajectory while keeping the task solvable, and use it to measure how well agents adapt.

Ask this paper

Key points
01

Method. AGNI reads a verified successful trajectory, extracts the assumptions it depends on (a writable directory, a package version, a command's behavior), injects a change that invalidates one of them, and keeps only variants where the old strategy fails but an adapted one still passes.

02

Adaptation gap. Average pass@1 drops from 84.1% to 53.4% on Endless Terminals, from 63.6% to 37.6% on TB-Lite and from 56.2% to 39.8% on Terminal-Bench 2 once the change is injected.

03

Base score does not predict adaptation. GPT-5.4-mini leads on base Endless Terminals tasks at 94.5% but falls to 61.9% on novel ones, below Kimi-2.6 (74.7%) and DeepSeek-V4-Flash (74.2%). Telling the agent what changed recovers 71.7% of the gap on average.

04

Where agents fail. Agents usually notice the change but do not diagnose it. GPT-5.5 and Grok-4.20 recognize the novelty at the same rate (about 80%), but GPT-5.5 reaches strategy revision in 77.5% of trajectories against 47.1% for Grok.

05

Training on novelty. GRPO on paired base and novel tasks cuts the adaptation gap from 45.8 to 28.6 points, and adding a meta-action prompt during training lifts held-out novel pass@1 to 37.9% while base pass@1 also rises.

Abstract

LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable agents must detect the change and adapt while pursuing the same goal. We study this adaptation capability through environmental novelty: a change that keeps the task objective fixed while invalidating an assumption underlying an otherwise successful trajectory. We introduce AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable. Across three terminal benchmarks, AGNI produces diverse novelties spanning resources, interfaces, constraints, and execution semantics. Evaluating multiple LLM agents reveals a substantial adaptation gap between base and novel tasks. Trajectory analysis suggests that agents often encounter evidence of the change but fail to diagnose its cause and revise their strategy. Finally, post-training for environmental novelty improves adaptation to held-out novel tasks while also improving performance on base tasks. Our results highlight a gap between task competence and adaptive capability and motivate environmental variation as a core dimension of agent training and evaluation.

Every Monday
Get next week’s papers.
Subscribe on Substack