🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
AgentBoard

AgentBoard

AgentBoard is a benchmark and open-source evaluation framework for analytically evaluating LLM agents beyond the usual pass/fail metrics.

1081Evaluation
Sleeper Agents

Sleeper Agents

Anthropic shows that LLMs can be trained to act deceptively under specific triggers and that current safety training techniques fail to remove this hidden behavior.

1082Safety
RAISE

RAISE

RAISE is an advanced agent architecture that adds a dual-memory system on top of a ReAct-style backbone to better support long-running conversational agents.

1083Memory
SeeAct (GPT-4V as Generalist Web Agent)

SeeAct (GPT-4V as Generalist Web Agent)

OSU researchers adapt GPT-4V into SeeAct, a generalist agent that operates live websites using vision + language planning.

1084Agents
How Code Empowers LLMs

How Code Empowers LLMs

A survey on why training LLMs with code data produces capabilities well beyond coding itself.

1085Agents
AppAgent

AppAgent

Introduces an LLM-based multimodal agent that operates real smartphone apps through touch actions and screenshots.

1086Multimodal
ReST Meets ReAct

ReST Meets ReAct

Proposes a ReAct-style agent that improves itself via reinforced self-training on its own reasoning traces.

1087Agents
Pearl

Pearl

Meta's Pearl is a production-ready reinforcement learning agent package designed for real-world deployment constraints.

1088Agents
GNoME

GNoME

DeepMind's Graph Networks for Materials Exploration (GNoME) is an AI system that discovered 2.2 million new crystal structures, including 380,000 thermodynamically stable ones.

1089Agents
Hitchhiker's Guide From CoT to Agents

Hitchhiker's Guide From CoT to Agents

A survey mapping the conceptual evolution from chain-of-thought reasoning to modern language-agent frameworks.

1090Agents
GAIA

GAIA

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

1091Agents
MedAgents

MedAgents

A collaborative multi-round framework for medical reasoning that uses role-playing LLM agents to improve accuracy and reasoning depth.

1092Reasoning
JARVIS-1

JARVIS-1

An open-world multimodal agent for Minecraft that combines perception, planning, and memory into a self-improving system.

1093Agents
LLMs Can Deceive Users (Trading Agent)

LLMs Can Deceive Users (Trading Agent)

Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

1094Agents
Branch-Solve-Merge (BSM)

Branch-Solve-Merge (BSM)

BSM decomposes LLM tasks into parallel sub-tasks via three LLM-programmed modules: branch, solve, and merge.

1095Agents
OpenAgents

OpenAgents

An open platform for running and hosting real-world language agents, including three distinct agent types.

1096Agents
AutoMix

AutoMix

AutoMix routes queries between LLMs of different sizes based on smaller-model confidence, saving cost without sacrificing quality.

1097Efficiency
Video Language Planning

Video Language Planning

Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

1098Multimodal
FireAct (Language Agent Fine-tuning)

FireAct (Language Agent Fine-tuning)

Explores fine-tuning LLMs specifically for language-agent use, demonstrating consistent gains over prompting alone.

1099Agents
Qwen

Qwen

Alibaba releases the Qwen family of open LLMs with strong tool-use and planning capabilities for language agents.

1100Agents
Compositional Foundation Models (HiP)

Compositional Foundation Models (HiP)

Proposes foundation models that compose multiple expert foundation models trained on different modalities to solve long-horizon goals.

1101Agents
OWL (LLMs for IT Operations)

OWL (LLMs for IT Operations)

Proposes OWL, an LLM specialized for IT operations through self-instruct fine-tuning on IT-specific tasks.

1102Evaluation
The Rise and Potential of LLM-Based Agents

The Rise and Potential of LLM-Based Agents

A comprehensive survey of LLM-based agents covering construction, capability, and societal implications.

1103Agents
Agents Library

Agents Library

An open-source library for building autonomous language agents with first-class support for planning, memory, tools, and multi-agent communication.

1104Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026