AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
AgentBoard
AgentBoard is a benchmark and open-source evaluation framework for analytically evaluating LLM agents beyond the usual pass/fail metrics.

Sleeper Agents
Anthropic shows that LLMs can be trained to act deceptively under specific triggers and that current safety training techniques fail to remove this hidden behavior.

RAISE
RAISE is an advanced agent architecture that adds a dual-memory system on top of a ReAct-style backbone to better support long-running conversational agents.

SeeAct (GPT-4V as Generalist Web Agent)
OSU researchers adapt GPT-4V into SeeAct, a generalist agent that operates live websites using vision + language planning.

How Code Empowers LLMs
A survey on why training LLMs with code data produces capabilities well beyond coding itself.

AppAgent
Introduces an LLM-based multimodal agent that operates real smartphone apps through touch actions and screenshots.

ReST Meets ReAct
Proposes a ReAct-style agent that improves itself via reinforced self-training on its own reasoning traces.

Pearl
Meta's Pearl is a production-ready reinforcement learning agent package designed for real-world deployment constraints.

GNoME
DeepMind's Graph Networks for Materials Exploration (GNoME) is an AI system that discovered 2.2 million new crystal structures, including 380,000 thermodynamically stable ones.

Hitchhiker's Guide From CoT to Agents
A survey mapping the conceptual evolution from chain-of-thought reasoning to modern language-agent frameworks.

GAIA
Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

MedAgents
A collaborative multi-round framework for medical reasoning that uses role-playing LLM agents to improve accuracy and reasoning depth.

JARVIS-1
An open-world multimodal agent for Minecraft that combines perception, planning, and memory into a self-improving system.

LLMs Can Deceive Users (Trading Agent)
Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

Branch-Solve-Merge (BSM)
BSM decomposes LLM tasks into parallel sub-tasks via three LLM-programmed modules: branch, solve, and merge.

OpenAgents
An open platform for running and hosting real-world language agents, including three distinct agent types.

AutoMix
AutoMix routes queries between LLMs of different sizes based on smaller-model confidence, saving cost without sacrificing quality.

Video Language Planning
Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

FireAct (Language Agent Fine-tuning)
Explores fine-tuning LLMs specifically for language-agent use, demonstrating consistent gains over prompting alone.

Qwen
Alibaba releases the Qwen family of open LLMs with strong tool-use and planning capabilities for language agents.

Compositional Foundation Models (HiP)
Proposes foundation models that compose multiple expert foundation models trained on different modalities to solve long-horizon goals.

OWL (LLMs for IT Operations)
Proposes OWL, an LLM specialized for IT operations through self-instruct fine-tuning on IT-specific tasks.

The Rise and Potential of LLM-Based Agents
A comprehensive survey of LLM-based agents covering construction, capability, and societal implications.

Agents Library
An open-source library for building autonomous language agents with first-class support for planning, memory, tools, and multi-agent communication.