Sapien: A Stateful Policy Engine for Autonomous AI Agents

Corinn Tiffany, Wen Zhang, Eugene Bagdasarian and Lillian Tsai at Google (with UMass Amherst) present Sapien, a policy engine that enforces task-specific policies on an agent's tool calls where what is allowed depends on what the agent has already done.
Ask this paper
Problem. Contextual security systems generate a per-task allowlist, but in multi-step tasks the valid next action depends on earlier actions and on data the agent has read, which a static allowlist cannot express.
Policy language. A Sapien policy is a regular expression over tool-call sequences, extended with stateful predicates, deferred policy generation for steps that cannot be specified up front, and scoped semantic checks.
Utility cost. Enforcing Sapien policies keeps task success within a few percent of an unconstrained agent.
Security result. Assuming the agent is fully hijacked, Sapien blocks 93 to 95% of attacks on AgentDojo and 62 to 85% on Toolathlon, about twice as many as tool allowlists on long-horizon tasks.
Abstract
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).