🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
Logical Chain-of-Thought (LogiCoT)

Logical Chain-of-Thought (LogiCoT)

A neurosymbolic framework that verifies and revises zero-shot CoT reasoning using symbolic-logic principles.

02Reasoning
AlphaMissense

AlphaMissense

DeepMind's AlphaMissense is an AI model that classifies missense genetic variants as pathogenic or benign at genome scale.

03Architecture
Chain-of-Verification (CoVe)

Chain-of-Verification (CoVe)

Meta's Chain-of-Verification adds a "deliberation" step where the LLM fact-checks its own draft before finalizing.

04Reasoning
Contrastive Decoding for Reasoning

Contrastive Decoding for Reasoning

Shows that contrastive decoding, a simple inference-time technique, substantially improves reasoning in large LLMs.

05Reasoning
LongLoRA

LongLoRA

An efficient LoRA-based fine-tuning recipe for extending LLM context windows without expensive full fine-tuning.

06Training
Struc-Bench (LLMs for Structured Data)

Struc-Bench (LLMs for Structured Data)

Studies how LLMs handle complex structured-data generation and proposes a structure-aware fine-tuning method.

07Training
LMSYS-Chat-1M

LMSYS-Chat-1M

LMSYS releases a large-scale dataset of 1 million real-world LLM conversations collected from the Vicuna demo and Chatbot Arena.

08Data
Language Modeling Is Compression

Language Modeling Is Compression

DeepMind empirically revisits the theoretical equivalence between prediction and compression, applied to modern LLMs.

09Efficiency
Compositional Foundation Models (HiP)

Compositional Foundation Models (HiP)

Proposes foundation models that compose multiple expert foundation models trained on different modalities to solve long-horizon goals.

10Agents
OWL (LLMs for IT Operations)

OWL (LLMs for IT Operations)

Proposes OWL, an LLM specialized for IT operations through self-instruct fine-tuning on IT-specific tasks.

11Evaluation
KOSMOS-2.5

KOSMOS-2.5

Microsoft's KOSMOS-2.5 is a multimodal model purpose-built for "machine reading" of text-intensive images.

12Multimodal
Textbooks Are All You Need II (phi-1.5)

Textbooks Are All You Need II (phi-1.5)

Microsoft's phi-1.5 demonstrates that a 1.3B model trained on "textbook-quality" synthetic data rivals much larger models on reasoning.

13Data
The Rise and Potential of LLM-Based Agents

The Rise and Potential of LLM-Based Agents

A comprehensive survey of LLM-based agents covering construction, capability, and societal implications.

14Agents
EvoDiff

EvoDiff

Microsoft's EvoDiff combines evolutionary-scale protein data with diffusion models for controllable protein generation in sequence space.

15Architecture
Rewindable Auto-regressive INference (RAIN)

Rewindable Auto-regressive INference (RAIN)

Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

16Safety
Robot Parkour Learning

Robot Parkour Learning

Stanford's Robot Parkour system learns end-to-end vision-based parkour policies that transfer to a quadrupedal robot.

17Robotics
Hallucination Survey (Early)

Hallucination Survey (Early)

Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

18Safety
Agents Library

Agents Library

An open-source library for building autonomous language agents with first-class support for planning, memory, tools, and multi-agent communication.

19Agents
Radiology-Llama 2

Radiology-Llama 2

A Llama 2-based LLM specialized for radiology report generation.

20Training
ChatDev (Communicative Agents for Software Development)

ChatDev (Communicative Agents for Software Development)

ChatDev is a virtual chat-powered software company where LLM agents take on roles in a waterfall-model dev process.

21Agents
MAmmoTH

MAmmoTH

An open-source LLM family specialized for general mathematical problem solving.

22Reasoning
Transformers as Support Vector Machines

Transformers as Support Vector Machines

A theoretical paper establishing a formal connection between self-attention optimization and hard-margin SVM problems.

23Training
RLAIF (Scaling RLHF with AI Feedback)

RLAIF (Scaling RLHF with AI Feedback)

Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

24Reinforcement Learning
GPT Solves Math Problems Without a Calculator

GPT Solves Math Problems Without a Calculator

Demonstrates that with sufficient training data, even a small language model can perform accurate multi-digit arithmetic.

25Reasoning
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026