Agents as Software: A Programming Languages Agenda for Agent Reliability

Shraddha Barke (Microsoft Research) and Adithya Murali (UW-Madison) argue in an Onward! 2026 essay that agents should be treated as programs whose behavior is spread across prompts, tools, memories and traces, and lay out how specifications, static analysis and runtime monitoring from programming languages research apply to them.
Ask this paper
Core claim. An agent's program is no longer one body of source code. It is distributed across prompts, tool definitions, memories, policies and orchestration logic, so ordinary testing and debugging cover only part of its behavior.
Specifications over traces. The authors propose writing agent requirements as properties over event logs and state: tool preconditions, permissions, information flow, provenance and freshness of observations, and resource limits, using temporal and first-order logics.
Static analysis. Agent configurations, prompts and tool schemas can be checked before deployment as static artifacts, for example validating that a workflow cannot reach a high-impact tool without a required prior step.
Runtime monitoring. The essay frames dynamic checks as questions a monitor must answer: is this tool call justified by evidence, is the observation still current, did the required step happen before a high-impact action, and what recovery runs after a violation.
Scope. The goal stated is to give probabilistic agents enough structure that their behavior can be reasoned about, controlled and repaired, without requiring them to behave like deterministic programs. It is an agenda paper with no experiments.
Abstract
AI agents increasingly resemble software systems: they call tools, remember facts, follow policies, delegate work, and take actions with real consequences. % Yet the ``program'' of an agent is scattered across prompts, tools, memories, workflows, and execution traces, making its behavior difficult to inspect through ordinary testing and debugging alone. % This essay argues that a programming-systems perspective offers a natural lens for making agents reliable. % We recast agents as programmable artifacts whose behavior can be specified over traces and state, checked before deployment, monitored during execution, and improved from observed failures. % The goal is not to make probabilistic agents behave like deterministic programs, but to give them enough structure that their behavior can be reasoned about, controlled, and repaired.