🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023Issue 182 · Sep 28 – Oct 4, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
This week · 10 papersView the full issue →
AlphaMissense

AlphaMissense

DeepMind's AlphaMissense is an AI model that classifies missense genetic variants as pathogenic or benign at genome scale.

02Architecture
Chain-of-Verification (CoVe)

Chain-of-Verification (CoVe)

Meta's Chain-of-Verification adds a "deliberation" step where the LLM fact-checks its own draft before finalizing.

03Reasoning
Contrastive Decoding for Reasoning

Contrastive Decoding for Reasoning

Shows that contrastive decoding, a simple inference-time technique, substantially improves reasoning in large LLMs.

04Reasoning
LongLoRA

LongLoRA

An efficient LoRA-based fine-tuning recipe for extending LLM context windows without expensive full fine-tuning.

05Training
Struc-Bench (LLMs for Structured Data)

Struc-Bench (LLMs for Structured Data)

Studies how LLMs handle complex structured-data generation and proposes a structure-aware fine-tuning method.

06Training
LMSYS-Chat-1M

LMSYS-Chat-1M

LMSYS releases a large-scale dataset of 1 million real-world LLM conversations collected from the Vicuna demo and Chatbot Arena.

07Data
Language Modeling Is Compression

Language Modeling Is Compression

DeepMind empirically revisits the theoretical equivalence between prediction and compression, applied to modern LLMs.

08Efficiency
Compositional Foundation Models (HiP)

Compositional Foundation Models (HiP)

Proposes foundation models that compose multiple expert foundation models trained on different modalities to solve long-horizon goals.

09Agents
OWL (LLMs for IT Operations)

OWL (LLMs for IT Operations)

Proposes OWL, an LLM specialized for IT operations through self-instruct fine-tuning on IT-specific tasks.

10Evaluation
KOSMOS-2.5

KOSMOS-2.5

Microsoft's KOSMOS-2.5 is a multimodal model purpose-built for "machine reading" of text-intensive images.

11Multimodal
Textbooks Are All You Need II (phi-1.5)

Textbooks Are All You Need II (phi-1.5)

Microsoft's phi-1.5 demonstrates that a 1.3B model trained on "textbook-quality" synthetic data rivals much larger models on reasoning.

12Data
The Rise and Potential of LLM-Based Agents

The Rise and Potential of LLM-Based Agents

A comprehensive survey of LLM-based agents covering construction, capability, and societal implications.

13Agents
EvoDiff

EvoDiff

Microsoft's EvoDiff combines evolutionary-scale protein data with diffusion models for controllable protein generation in sequence space.

14Architecture
Rewindable Auto-regressive INference (RAIN)

Rewindable Auto-regressive INference (RAIN)

Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

15Safety
Robot Parkour Learning

Robot Parkour Learning

Stanford's Robot Parkour system learns end-to-end vision-based parkour policies that transfer to a quadrupedal robot.

16Robotics
Hallucination Survey (Early)

Hallucination Survey (Early)

Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

17Safety
Agents Library

Agents Library

An open-source library for building autonomous language agents with first-class support for planning, memory, tools, and multi-agent communication.

18Agents
Radiology-Llama 2

Radiology-Llama 2

A Llama 2-based LLM specialized for radiology report generation.

19Training
ChatDev (Communicative Agents for Software Development)

ChatDev (Communicative Agents for Software Development)

ChatDev is a virtual chat-powered software company where LLM agents take on roles in a waterfall-model dev process.

20Agents
MAmmoTH

MAmmoTH

An open-source LLM family specialized for general mathematical problem solving.

21Reasoning
Transformers as Support Vector Machines

Transformers as Support Vector Machines

A theoretical paper establishing a formal connection between self-attention optimization and hard-margin SVM problems.

22Training
RLAIF (Scaling RLHF with AI Feedback)

RLAIF (Scaling RLHF with AI Feedback)

Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

23Reinforcement Learning
GPT Solves Math Problems Without a Calculator

GPT Solves Math Problems Without a Calculator

Demonstrates that with sufficient training data, even a small language model can perform accurate multi-digit arithmetic.

24Reasoning
OPRO (LLMs as Optimizers)

OPRO (LLMs as Optimizers)

DeepMind's OPRO uses LLMs as general-purpose optimizers over natural-language-described problems.

25Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026