🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
Med-Flamingo

Med-Flamingo

Stanford's Med-Flamingo is a multimodal medical model supporting in-context learning for few-shot medical visual QA.

02Multimodal
ToolLLM

ToolLLM

Tsinghua's ToolLLM enables LLMs to interact with 16,000+ real-world APIs through a comprehensive framework for tool-using LLMs.

03Agents
Skeleton-of-Thought (SoT)

Skeleton-of-Thought (SoT)

Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.

04Reasoning
MetaGPT

MetaGPT

MetaGPT is a multi-agent framework that encodes standardized operating procedures (SOPs) for complex problem solving.

05Agents
OpenFlamingo

OpenFlamingo

An open-source family of autoregressive vision-language models spanning 3B to 9B parameters.

06Multimodal
The Hydra Effect

The Hydra Effect

DeepMind shows that language models exhibit self-repairing behavior when attention heads are ablated.

07Safety
Self-Check

Self-Check

Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

08Reasoning
Dynalang (Agents Model the World with Language)

Dynalang (Agents Model the World with Language)

UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

09Agents
AutoRobotics-Zero

AutoRobotics-Zero

Discovers zero-shot adaptable robot policies from scratch, including the automatic discovery of Python control code.

10Robotics
Universal Adversarial LLM Attacks

Universal Adversarial LLM Attacks

Finds universal and transferable adversarial attacks that cause aligned models like ChatGPT and Bard to generate objectionable behaviors.

11Safety
RT-2

RT-2

Google DeepMind's end-to-end vision-language-action model that learns from both web and robotics data to control robots.

12Robotics
Med-PaLM Multimodal

Med-PaLM Multimodal

Introduces a generalist biomedical AI system and a new multimodal biomedical benchmark with 14 tasks.

13Multimodal
Tracking Anything in High Quality

Tracking Anything in High Quality

A framework for high-quality tracking-anything in videos combining segmentation and refinement.

14Multimodal
Foundation Models in Vision

Foundation Models in Vision

A comprehensive survey on foundational models for computer vision and their open research directions.

15Multimodal
L-Eval

L-Eval

A standardized evaluation suite for long-context language models.

16Evaluation
LoraHub

LoraHub

Enables efficient cross-task generalization via dynamic LoRA composition.

17Training
Survey of Aligned LLMs

Survey of Aligned LLMs

A comprehensive overview of alignment approaches covering data, training, and evaluation.

18Safety
WavJourney

WavJourney

Leverages LLMs to orchestrate audio generation models for compositional storytelling.

19Multimodal
FacTool

FacTool

A task- and domain-agnostic framework for factuality detection of LLM-generated text.

20Evaluation
Llama 2

Llama 2

Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

21Training
How is ChatGPT's Behavior Changing Over Time?

How is ChatGPT's Behavior Changing Over Time?

Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

22Evaluation
FlashAttention-2

FlashAttention-2

Tri Dao's follow-up to FlashAttention, dramatically improving attention throughput on modern GPUs.

23Efficiency
Measuring Faithfulness in Chain-of-Thought Reasoning

Measuring Faithfulness in Chain-of-Thought Reasoning

Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.

24Reasoning
Generative TV & Showrunner Agents

Generative TV & Showrunner Agents

Fable Studio's approach to generate episodic TV content using LLMs and multi-agent simulation.

25Agents
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026