🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
623 papers · AgentsClear filters →
How Code Empowers LLMs

How Code Empowers LLMs

A survey on why training LLMs with code data produces capabilities well beyond coding itself.

553Agents
CogAgent

CogAgent

Tsinghua's CogAgent is an 18B-parameter visual-language model purpose-built for GUI understanding and navigation, with unusually high input resolution.

554Evaluation
From Gemini to Q-Star

From Gemini to Q-Star

A 300+-paper survey mapping the state of Generative AI and the research frontiers that followed the Gemini + rumored Q* news cycle.

555Multimodal
Survey of Reasoning with Foundation Models

Survey of Reasoning with Foundation Models

A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

556Reasoning
AppAgent

AppAgent

Introduces an LLM-based multimodal agent that operates real smartphone apps through touch actions and screenshots.

557Multimodal
ReST Meets ReAct

ReST Meets ReAct

Proposes a ReAct-style agent that improves itself via reinforced self-training on its own reasoning traces.

558Agents
RAG for LLMs

RAG for LLMs

A broad survey of Retrieval-Augmented Generation research, organizing the rapidly growing literature into a coherent map.

559Retrieval
Pearl

Pearl

Meta's Pearl is a production-ready reinforcement learning agent package designed for real-world deployment constraints.

560Agents
GNoME

GNoME

DeepMind's Graph Networks for Materials Exploration (GNoME) is an AI system that discovered 2.2 million new crystal structures, including 380,000 thermodynamically stable ones.

561Agents
Hitchhiker's Guide From CoT to Agents

Hitchhiker's Guide From CoT to Agents

A survey mapping the conceptual evolution from chain-of-thought reasoning to modern language-agent frameworks.

562Agents
GAIA

GAIA

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

563Agents
MedAgents

MedAgents

A collaborative multi-round framework for medical reasoning that uses role-playing LLM agents to improve accuracy and reasoning depth.

564Reasoning
JARVIS-1

JARVIS-1

An open-world multimodal agent for Minecraft that combines perception, planning, and memory into a self-improving system.

565Agents
LLMs Can Deceive Users (Trading Agent)

LLMs Can Deceive Users (Trading Agent)

Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

566Agents
On the Road with GPT-4V

On the Road with GPT-4V

An exhaustive evaluation of GPT-4V applied to autonomous driving scenarios.

567Evaluation
Managing AI Risks (Bengio, Hinton, et al.)

Managing AI Risks (Bengio, Hinton, et al.)

A high-profile position paper by leading AI researchers laying out risks from upcoming advanced AI systems.

568Safety
Branch-Solve-Merge (BSM)

Branch-Solve-Merge (BSM)

BSM decomposes LLM tasks into parallel sub-tasks via three LLM-programmed modules: branch, solve, and merge.

569Agents
LLMs for Software Engineering

LLMs for Software Engineering

A comprehensive survey of LLMs for software engineering covering models, tasks, evaluation, and open challenges.

570Code
OpenAgents

OpenAgents

An open platform for running and hosting real-world language agents, including three distinct agent types.

571Agents
AutoMix

AutoMix

AutoMix routes queries between LLMs of different sizes based on smaller-model confidence, saving cost without sacrificing quality.

572Efficiency
Video Language Planning

Video Language Planning

Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

573Multimodal
UniSim (Universal Simulator)

UniSim (Universal Simulator)

Google's UniSim learns a universal generative simulator of real-world interactions from diverse video + action data.

574Robotics
MemWalker

MemWalker

MemWalker treats the LLM as an interactive agent that traverses a tree-structured summary of long text.

575Memory
FireAct (Language Agent Fine-tuning)

FireAct (Language Agent Fine-tuning)

Explores fine-tuning LLMs specifically for language-agent use, demonstrating consistent gains over prompting alone.

576Agents
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026