🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023Issue 182 · Sep 28 – Oct 4, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
This week · 10 papersView the full issue →
ToolLLM

ToolLLM

Tsinghua's ToolLLM enables LLMs to interact with 16,000+ real-world APIs through a comprehensive framework for tool-using LLMs.

02Agents
Skeleton-of-Thought (SoT)

Skeleton-of-Thought (SoT)

Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.

03Reasoning
MetaGPT

MetaGPT

MetaGPT is a multi-agent framework that encodes standardized operating procedures (SOPs) for complex problem solving.

04Agents
OpenFlamingo

OpenFlamingo

An open-source family of autoregressive vision-language models spanning 3B to 9B parameters.

05Multimodal
The Hydra Effect

The Hydra Effect

DeepMind shows that language models exhibit self-repairing behavior when attention heads are ablated.

06Safety
Self-Check

Self-Check

Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

07Reasoning
Dynalang (Agents Model the World with Language)

Dynalang (Agents Model the World with Language)

UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

08Agents
AutoRobotics-Zero

AutoRobotics-Zero

Discovers zero-shot adaptable robot policies from scratch, including the automatic discovery of Python control code.

09Robotics
Universal Adversarial LLM Attacks

Universal Adversarial LLM Attacks

Finds universal and transferable adversarial attacks that cause aligned models like ChatGPT and Bard to generate objectionable behaviors.

10Safety
RT-2

RT-2

Google DeepMind's end-to-end vision-language-action model that learns from both web and robotics data to control robots.

11Robotics
Med-PaLM Multimodal

Med-PaLM Multimodal

Introduces a generalist biomedical AI system and a new multimodal biomedical benchmark with 14 tasks.

12Multimodal
Tracking Anything in High Quality

Tracking Anything in High Quality

A framework for high-quality tracking-anything in videos combining segmentation and refinement.

13Multimodal
Foundation Models in Vision

Foundation Models in Vision

A comprehensive survey on foundational models for computer vision and their open research directions.

14Multimodal
L-Eval

L-Eval

A standardized evaluation suite for long-context language models.

15Evaluation
LoraHub

LoraHub

Enables efficient cross-task generalization via dynamic LoRA composition.

16Training
Survey of Aligned LLMs

Survey of Aligned LLMs

A comprehensive overview of alignment approaches covering data, training, and evaluation.

17Safety
WavJourney

WavJourney

Leverages LLMs to orchestrate audio generation models for compositional storytelling.

18Multimodal
FacTool

FacTool

A task- and domain-agnostic framework for factuality detection of LLM-generated text.

19Evaluation
Llama 2

Llama 2

Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

20Training
How is ChatGPT's Behavior Changing Over Time?

How is ChatGPT's Behavior Changing Over Time?

Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

21Evaluation
FlashAttention-2

FlashAttention-2

Tri Dao's follow-up to FlashAttention, dramatically improving attention throughput on modern GPUs.

22Efficiency
Measuring Faithfulness in Chain-of-Thought Reasoning

Measuring Faithfulness in Chain-of-Thought Reasoning

Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.

23Reasoning
Generative TV & Showrunner Agents

Generative TV & Showrunner Agents

Fable Studio's approach to generate episodic TV content using LLMs and multi-agent simulation.

24Agents
Challenges & Application of LLMs

Challenges & Application of LLMs

A comprehensive enumeration of open challenges and application domains for LLMs.

25Safety
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026