Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety

Charlie Summers, Oliver Kennedy (Buffalo), Eugene Wu and colleagues at Columbia propose Environment Steering: model the agent's and harness's execution state as database tables, track record-level data flow, and when a declarative policy is violated, send the agent specific feedback so it can recover safely.
Ask this paper
Approach. Instead of filtering prompts or asking an LLM judge, the environment checks data provenance before each action commits, using data flow control techniques from databases.
Steering, not only blocking. When a violation is detected, the agent receives policy- and context-specific feedback that points it toward a safe alternative, so tasks can still complete.
AgentDyn result. With Qwen3 235B, it reaches 61% task success with 0% attack success, compared against 9 reference defenses, improving the safety-utility frontier.
Across four safety benchmarks. It reduces attack success while raising task success by up to 12.7% over no defense.
Venue. REALM Workshop at EMNLP 2026.
Abstract
LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without helping the agent recover. We argue that the execution environment should instead enforce safety as the agent runs and steer it toward safe alternatives when violations occur---we call this Environment Steering. We implement this by modeling the agent and harness execution state as database tables, track the record-level data flows, and check these data flows against declarative policies during runtime. When violations are detected, policy- and context-specific feedback steers the agent toward safe trajectories. On AgentDyn, this enables the agent to improve task success rate over no-defense while achieving 0% attack success rate.