🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 18, 2026
Reinforcement Learning

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

First page
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
The curator’s take

Juzheng Zhang and colleagues at AWS AI Labs and the University of Maryland introduce ActObs, which applies SFT loss to the environment observation tokens already present in agent trajectories rather than only to agent-authored action tokens.

Ask this paper

Key points
01

No new data, parameters, tokens or forward passes. The observation tokens are already in each trajectory; only the loss mask changes, so the comparison against action-only SFT is clean.

02

The two are equal after SFT and diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget on Terminal-Bench 2.0.

03

At 8B it trades pass@1 for pass@k. ActObs gives up some pass@1 reliability for +3.4 points at pass@16 and solves more distinct tasks.

04

The advantage transfers across domains. On aider-polyglot code editing, unseen during both SFT and RL, ActObs gains 4.2 points at pass@1 at 4B.

05

The gradient analysis explains the mechanism. Action and observation gradients become orthogonal quickly under joint supervision, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model, so the policy enters RL with less exploration capacity.

Abstract

Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.

Every Monday
Get next week’s papers.
Subscribe on Substack