🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023Issue 182 · Sep 28 – Oct 4, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
This week · 10 papersView the full issue →
RECOMP (Retrieval-Augmented LMs with Compressors)

RECOMP (Retrieval-Augmented LMs with Compressors)

Proposes two compression approaches to shrink retrieved documents before in-context use.

02Retrieval
InstructRetro

InstructRetro

NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.

03Training
MemWalker

MemWalker

MemWalker treats the LLM as an interactive agent that traverses a tree-structured summary of long text.

04Memory
FireAct (Language Agent Fine-tuning)

FireAct (Language Agent Fine-tuning)

Explores fine-tuning LLMs specifically for language-agent use, demonstrating consistent gains over prompting alone.

05Agents
LLMs Represent Space and Time

LLMs Represent Space and Time

MIT researchers find that LLMs internally encode linear representations of space and time across multiple scales.

06Safety
Retrieval Meets Long-Context LLMs

Retrieval Meets Long-Context LLMs

NVIDIA's study comparing RAG and long-context LLMs, with the punchline that the two are complementary rather than substitutes.

07Retrieval
StreamingLLM

StreamingLLM

MIT's StreamingLLM enables efficient streaming inference by preserving "attention sinks" - early-sequence tokens that most attention mass flows to.

08Memory
Neural Developmental Programs (NDPs)

Neural Developmental Programs (NDPs)

Proposes neural networks that self-assemble through a developmental process inspired by biological embryonic development.

09Architecture
The Dawn of LMMs (GPT-4V Deep Dive)

The Dawn of LMMs (GPT-4V Deep Dive)

Microsoft's exhaustive 166-page analysis of GPT-4V's capabilities and limitations.

10Multimodal
Training LLMs with Pause Tokens

Training LLMs with Pause Tokens

CMU shows that adding a learnable `<pause>` token during both pretraining and fine-tuning gives the model extra "thinking time" and improves reasoning.

11Reasoning
Self-Taught Optimizer (STOP)

Self-Taught Optimizer (STOP)

Proposes recursively self-improving code generation where an LLM-scaffolded program improves itself.

12Code
RA-DIT (Retrieval-Augmented Dual Instruction Tuning)

RA-DIT (Retrieval-Augmented Dual Instruction Tuning)

Meta's RA-DIT is a lightweight recipe that retrofits LLMs with retrieval capabilities through dual fine-tuning.

13Retrieval
KOSMOS-G

KOSMOS-G

Microsoft's KOSMOS-G extends zero-shot image generation to multi-image vision-language input.

14Multimodal
Analogical Prompting

Analogical Prompting

Google's Analogical Prompting guides LLM reasoning by having the model self-generate relevant exemplars on the fly.

15Reasoning
The Reversal Curse

The Reversal Curse

Finds that LLMs trained on "A is B" fail to generalize to "B is A" - a surprisingly deep failure of learning.

16Training
Effective Long-Context Scaling (Meta)

Effective Long-Context Scaling (Meta)

Meta proposes a 70B long-context LLM that surpasses GPT-3.5-turbo-16k on long-context benchmarks.

17Memory
Graph Neural Prompting (GNP)

Graph Neural Prompting (GNP)

A plug-and-play method that injects knowledge-graph information into frozen pretrained LLMs.

18Training
Vision Transformers Need Registers

Vision Transformers Need Registers

Meta researchers identify artifact tokens in ViT feature maps and propose a trivial fix: add dedicated register tokens.

19Architecture
Boolformer

Boolformer

The first Transformer trained to perform end-to-end symbolic regression of Boolean functions.

20Reasoning
LLaVA-RLHF

LLaVA-RLHF

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

21Reinforcement Learning
LLM Alignment Survey

LLM Alignment Survey

A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.

22Safety
Qwen

Qwen

Alibaba releases the Qwen family of open LLMs with strong tool-use and planning capabilities for language agents.

23Agents
MentaLLaMA

MentaLLaMA

An open-source LLM family specialized for interpretable mental-health analysis on social media.

24Safety
Logical Chain-of-Thought (LogiCoT)

Logical Chain-of-Thought (LogiCoT)

A neurosymbolic framework that verifies and revises zero-shot CoT reasoning using symbolic-logic principles.

25Reasoning
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026