🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 7 – Sep 7, 2026
Agents

From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

First page
From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments
The curator’s take

Linsen Zhu and Mengqing Cai review the agentic-AI literature through 31 August 2026 and separate three things the field routinely conflates: model competence, harness integration, and the authority a deployment actually grants an agent.

Ask this paper

Key points
01

Three axes instead of one autonomy scale. The survey organizes evidence along delegated authority, temporal persistence, and environmental coupling, and insists on separating model, harness, and environment when attributing a result.

02

Action scope has grown faster than reliability. Across the papers examined, expansion of action interfaces is documented far better than robust task completion, error recovery, authorization, or independent verification.

03

Protocols are interoperability, not trust. MCP and Agent2Agent standardize how agents connect but supply no evidence that delegation is trustworthy; multi-agent organization adds specialization together with cost and correlated failure.

04

Justified delegation as the design rule. The authors propose expanding an agent's action scope only where there is evidence of provenance, bounded authority, failure detection, safe recovery, and calibrated human control.

05

Research agenda. Coupled model-plus-harness evaluation, capability-based permissions, durable state, cross-agent accountability, and staged physical validation.

Abstract

Large language models become consequential agents when surrounding systems let outputs change external state. Models now call tools, operate interfaces, delegate work, retain state, inhabit generated worlds, and control robots or laboratory equipment. Such advances are often narrated as one march toward autonomy, conflating model competence, system integration, persistence, and safe authority. This critical review synthesizes primary research and official technical specifications available by 31 August 2026. We organize the evidence along delegated authority, temporal persistence, and environmental coupling, while separating model, harness, and environment. Within the evidence examined, action-interface expansion is documented more convincingly than robust completion, recovery, authorization, or independent verification. Model Context Protocol and Agent2Agent improve interoperability but do not establish trustworthy delegation; multi-agent organization adds specialization alongside cost and correlated failure. Persistent simulations and world models support training and planning but do not themselves demonstrate agency; robotics and self-driving laboratories establish bounded feasibility rather than unattended open-world reliability. We propose justified delegation as an analytical and normative heuristic, not an observed law or certified score: expand action scope only where evidence supports provenance, bounded authority, failure detection, safe recovery, and calibrated human control. This framing yields a research agenda for coupled model-harness evaluation, capability-based permissions, durable state, cross-agent accountability, and staged physical validation.

Every Monday
Get next week’s papers.
Subscribe on Substack