🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents

Deep Research Agents

Free while signed in. Answers cite the passages they came from.

First page
Deep Research Agents
The curator’s take

Provides the most comprehensive survey to date of Deep Research (DR) agents, LLM-powered systems built for autonomous, multi-step informational research. The paper defines DR agents as AI systems that tightly integrate dynamic reasoning, adaptive long-horizon planning, tool use, retrieval, and structured report generation. It establishes a taxonomy of DR architectures, evaluates recent advances, and outlines the limitations of current systems and benchmarks.

Key points
01

The authors differentiate DR agents from classical RAG or tool-use pipelines by emphasizing their autonomy, continual reasoning, and adaptive planning. Static workflows like AI Scientist and AgentRxiv rely on rigid pipelines, while dynamic DR agents like OpenAI DR and Gemini DR can replan and adapt in response to intermediate results.

02

A key contribution is the classification of DR agents along three axes: (1) static vs dynamic workflows, (2) planning-only vs intent-planning strategies, and (3) single-agent vs multi-agent systems. For example, Grok DeepSearch uses a single-agent loop with sparse attention and dynamic tool use, while OpenManus and OWL adopt multi-agent orchestration with role specialization.

03

DR agents use both API-based and browser-based search methods. The former (e.g., arXiv, Google Search) is fast and structured, while the latter (e.g., Chromium-based agents like Manus and DeepResearcher) handles dynamic content and complex UI interactions, albeit at higher latency and fragility.

04

Most agents now integrate tool-use modules such as code execution, data analytics, and multimodal reasoning. Advanced agents like AutoGLM Rumination even feature computer use, enabling direct API calls and platform interactions (e.g., CNKI, WeChat), effectively bridging inference and execution.

05

Benchmarking remains immature. The paper catalogs agent performance across QA (e.g., HotpotQA, GPQA, HLE) and execution benchmarks (e.g., GAIA, SWE-Bench), but notes that many evaluations fail to test retrieval rigor or long-form synthesis. Benchmarks like BrowseComp and HLE are highlighted as essential next steps for grounding evaluation in open-ended, time-sensitive tasks.

06

Future challenges include integrating private or API-gated data sources, enabling parallel DAG-style execution, improving fact-checking via reflective reasoning, and creating structured benchmarks for multimodal report generation. The authors advocate for AI-native browsers, hierarchical reinforcement learning, and dynamic workflow mutation as enablers of true agentic autonomy.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack