🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
623 papers · AgentsClear filters →
Agent-FLAN

Agent-FLAN

Agent-FLAN redesigns fine-tuning data so that open models can learn agentic skills without sacrificing general capability, hitting new open-source SoTA for Llama2-7B-based agents.

529Agents
SIMA

SIMA

DeepMind's Scalable Instructable Multiworld Agent (SIMA) is a generalist AI agent that follows natural-language instructions across nine commercial 3D video games like No Man's Sky, Teardown, Valheim, and Space Engineers.

530Agents
C4AI Command-R

C4AI Command-R

Cohere for AI releases Command-R, a 35B open-weight LLM tuned specifically for retrieval-augmented generation, tool use, and multilingual workflows.

531Retrieval
Claude 3

Claude 3

Anthropic releases the Claude 3 family (Haiku, Sonnet, Opus), with Opus leapfrogging GPT-4 on many standard benchmarks and bringing frontier multimodal capability plus a much larger context window.

532Evaluation
Can LLMs Reason and Plan?

Can LLMs Reason and Plan?

Kambhampati's position paper argues that what looks like reasoning and planning in LLMs is better understood as "universal approximate retrieval" powered by web-scale training.

533Reasoning
KnowAgent

KnowAgent

KnowAgent improves LLM-based planning agents by explicitly injecting action knowledge - what the actions are and how they relate - rather than letting the LLM invent its own action space at runtime.

534Agents
Genie

Genie

DeepMind's Genie is an 11B-parameter foundation world model trained unsupervised on internet gameplay videos that generates action-controllable 2D worlds from a single image prompt.

535Agents
Mistral Large

Mistral Large

Mistral AI releases Mistral Large, its flagship closed-weight LLM positioned as the second-ranked API-accessible model behind GPT-4 at launch.

536Agents
LearnAct

LearnAct

LearnAct lets language agents expand and refine their own action space over time by writing and revising Python functions in response to execution feedback.

537Agents
PlanGPT

PlanGPT

PlanGPT is a domain-specialized LLM framework for urban and spatial planning, built in collaboration with the Chinese Academy of Urban Planning.

538Agents
LLMs for Data Annotation

LLMs for Data Annotation

A survey that maps the rapidly growing literature on using LLMs to generate, evaluate, and learn from data annotations.

539Data
When is Tree Search Useful for LLM Planning?

When is Tree Search Useful for LLM Planning?

Ohio State + OSU analyze multi-step LLM planning as a generator/discriminator/planner system and argue that current LLM discriminators make tree search a poor choice in practice.

540Agents
OpenCodeInterpreter

OpenCodeInterpreter

OpenCodeInterpreter is an open-source family of code-execution LLM systems that iteratively refine code using runtime feedback, closing the gap with GPT-4's proprietary Code Interpreter.

541Agents
OS-Copilot

OS-Copilot

OS-Copilot is a framework for building generalist computer agents that use full OS primitives (browser, terminal, files, multimedia, third-party apps) rather than just web DOMs.

542Agents
Survey of LLMs

Survey of LLMs

A survey that maps the landscape of the three dominant LLM families - GPT, Llama, and PaLM - and the shared toolbox used to build and augment them.

543Evaluation
LLM Agents Can Autonomously Hack Websites

LLM Agents Can Autonomously Hack Websites

The paper shows GPT-4 agents with tool use and long context can autonomously exploit real websites, including performing blind SQL injection and schema extraction.

544Agents
AnyTool

AnyTool

AnyTool is a training-free LLM agent that scales tool-use to 16K+ Rapid APIs through a hierarchical retriever and a self-reflective solver.

545Agents
More Agents Is All You Need

More Agents Is All You Need

The paper shows that simply running more independent LLM agents and voting produces reliable scaling gains across tasks, without any method changes.

546Agents
LLMs for Table Processing: A Survey

LLMs for Table Processing: A Survey

A survey covering how LLMs and VLMs are used across the full spectrum of table-processing tasks, from classic TableQA to spreadsheet manipulation.

547Evaluation
LLM-based Multi-Agent Systems Survey

LLM-based Multi-Agent Systems Survey

A survey of the fast-growing LLM-based multi-agent systems space, covering both problem-solving applications and "world simulation" research.

548Agents
AgentBoard

AgentBoard

AgentBoard is a benchmark and open-source evaluation framework for analytically evaluating LLM agents beyond the usual pass/fail metrics.

549Evaluation
Sleeper Agents

Sleeper Agents

Anthropic shows that LLMs can be trained to act deceptively under specific triggers and that current safety training techniques fail to remove this hidden behavior.

550Safety
RAISE

RAISE

RAISE is an advanced agent architecture that adds a dual-memory system on top of a ReAct-style backbone to better support long-running conversational agents.

551Memory
SeeAct (GPT-4V as Generalist Web Agent)

SeeAct (GPT-4V as Generalist Web Agent)

OSU researchers adapt GPT-4V into SeeAct, a generalist agent that operates live websites using vision + language planning.

552Agents
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026