🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
Seamless

Seamless

Meta's Seamless is a family of models for end-to-end expressive, streaming cross-lingual speech communication.

02Multimodal
MEDITRON-70B

MEDITRON-70B

EPFL's MEDITRON is an open-source family of medical LLMs at 7B and 70B parameters, continually pretrained on curated medical corpora.

03Training
Medprompt

Medprompt

Microsoft researchers show that careful prompt engineering can push general-purpose GPT-4 to state-of-the-art on medical benchmarks, no domain fine-tuning required.

04Reasoning
UniIR

UniIR

UniIR is a unified instruction-guided multimodal retriever that handles eight retrieval tasks across modalities with a single model.

05Multimodal
Safe Deployment of Generative AI (Nature)

Safe Deployment of Generative AI (Nature)

A Nature correspondence arguing that medical professionals - not commercial interests - must drive the development and deployment of generative AI in medicine.

06Safety
Dobb-E

Dobb-E

NYU's Dobb-E is an affordable household-manipulation robot that learns new tasks with just 5 minutes of user demonstrations.

07Robotics
Translatotron 3

Translatotron 3

Google's Translatotron 3 performs speech-to-speech translation using only monolingual data - no parallel corpora required.

08Multimodal
System 2 Attention (S2A)

System 2 Attention (S2A)

Meta's S2A uses the LLM's own reasoning to decide what context actually matters, regenerating a clean prompt before the final response step.

09Reasoning
Advancing Long-Context LLMs

Advancing Long-Context LLMs

A survey of methodologies for improving Transformer long-context capability across pretraining, fine-tuning, and inference stages.

10Memory
Parallel Speculative Sampling

Parallel Speculative Sampling

Amazon researchers propose a parallel variant of speculative sampling that achieves significant LLM inference speedups with minimal extra parameters.

11Efficiency
Mirasol3B

Mirasol3B

Google's Mirasol3B is a multimodal model that decouples modalities into focused autoregressive components rather than forcing a single fused stream.

12Multimodal
Teaching Small LMs to Reason

Teaching Small LMs to Reason

An approach that teaches smaller language models to explicitly select among reasoning techniques for each problem.

13Reasoning
GPQA

GPQA

A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.

14Reasoning
Hitchhiker's Guide From CoT to Agents

Hitchhiker's Guide From CoT to Agents

A survey mapping the conceptual evolution from chain-of-thought reasoning to modern language-agent frameworks.

15Agents
GAIA

GAIA

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

16Agents
MedAgents

MedAgents

A collaborative multi-round framework for medical reasoning that uses role-playing LLM agents to improve accuracy and reasoning depth.

17Reasoning
TÜLU 2

TÜLU 2

Allen AI's TÜLU 2 is a suite of improved open instruction-tuned LLMs and an accompanying study of adaptation best practices.

18Training
Emu Video and Emu Edit

Emu Video and Emu Edit

Meta releases Emu Video and Emu Edit, a pair of diffusion models targeting controlled text-to-video generation and instruction-based image editing.

19Multimodal
Chain-of-Note (CoN)

Chain-of-Note (CoN)

Tencent's Chain-of-Note adds an explicit note-taking step to RAG so the model can evaluate retrieved evidence before answering.

20Retrieval
LLMs for Scientific Discovery

LLMs for Scientific Discovery

A broad evaluation of GPT-4 across scientific disciplines including drug discovery, biology, and computational chemistry.

21Evaluation
Fine-Tuning LLMs for Factuality

Fine-Tuning LLMs for Factuality

Stanford fine-tunes LLMs for factuality without any human labels by using automatically generated preference signals.

22Training
Contrastive Chain-of-Thought

Contrastive Chain-of-Thought

Proposes contrastive CoT prompting where models see both valid *and* invalid reasoning demonstrations to reduce reasoning errors.

23Reasoning
Survey on Language Models for Code

Survey on Language Models for Code

A comprehensive survey of LLMs for code covering 50+ models, 30+ evaluation tasks, and 500 related works.

24Code
JARVIS-1

JARVIS-1

An open-world multimodal agent for Minecraft that combines perception, planning, and memory into a self-improving system.

25Agents
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026