🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
GRIT

GRIT

GRIT (Generative Representational Instruction Tuning) trains a single LLM to handle both generative and embedding tasks, switching behavior based on instructions.

02Retrieval
LoRA+

LoRA+

LoRA+ is a minimal one-line change to LoRA: use different learning rates for the down-projection (A) and up-projection (B) matrices to restore feature learning at large width.

03Training
Back to Basics: Revisiting REINFORCE in RLHF

Back to Basics: Revisiting REINFORCE in RLHF

Cohere researchers argue that PPO is overkill for RLHF and that a simpler REINFORCE-style estimator works better in practice.

04Reinforcement Learning
Recurrent Memory Finds What LLMs Miss

Recurrent Memory Finds What LLMs Miss

Introduces BABILong, a new long-context benchmark, and shows that transformers with recurrent memory can handle sequences far beyond vanilla LLMs.

05Memory
When is Tree Search Useful for LLM Planning?

When is Tree Search Useful for LLM Planning?

Ohio State + OSU analyze multi-step LLM planning as a generator/discriminator/planner system and argue that current LLM discriminators make tree search a poor choice in practice.

06Agents
Chain-of-Thought Reasoning Without Prompting

Chain-of-Thought Reasoning Without Prompting

DeepMind shows that LLMs often *already* emit chain-of-thought reasoning in alternative decoding paths, and that selecting those paths via confidence lifts reasoning accuracy with no prompt engineering.

07Reasoning
OpenCodeInterpreter

OpenCodeInterpreter

OpenCodeInterpreter is an open-source family of code-execution LLM systems that iteratively refine code using runtime feedback, closing the gap with GPT-4's proprietary Code Interpreter.

08Code
Sora

Sora

OpenAI unveils Sora, a text-to-video diffusion-transformer that generates coherent, minute-long 1080p videos from natural-language prompts.

09Multimodal
Gemini 1.5

Gemini 1.5

Google DeepMind's Gemini 1.5 is a multimodal MoE LLM that scales context to 1M tokens (10M in research settings) while matching or surpassing Gemini 1.0 Ultra on standard benchmarks.

10Memory
V-JEPA

V-JEPA

Meta's V-JEPA learns visual representations by predicting features in masked video regions, without pretrained image encoders, text, negatives, or reconstruction.

11Training
Large World Model (LWM)

Large World Model (LWM)

UC Berkeley's LWM is an open 7B multimodal model trained on long videos and books that handles context windows up to 1M tokens via RingAttention.

12Memory
The Boundary of Neural Network Trainability is Fractal

The Boundary of Neural Network Trainability is Fractal

Sohl-Dickstein finds that the boundary between trainable and untrainable hyperparameter configurations looks like a Mandelbrot-style fractal across many architectures.

13Training
OS-Copilot

OS-Copilot

OS-Copilot is a framework for building generalist computer agents that use full OS primitives (browser, terminal, files, multimedia, third-party apps) rather than just web DOMs.

14Agents
TestGen-LLM

TestGen-LLM

Meta's TestGen-LLM uses LLMs to improve existing human-written tests - augmenting coverage rather than generating tests from scratch - while rigorously filtering LLM output for quality.

15Code
ChemLLM

ChemLLM

ChemLLM is a chemistry-specialized LLM with a matched dataset (ChemData) and benchmark (ChemBench) for evaluating chemistry-specific capability.

16Evaluation
Survey of LLMs

Survey of LLMs

A survey that maps the landscape of the three dominant LLM families - GPT, Llama, and PaLM - and the shared toolbox used to build and augment them.

17Evaluation
LLM Agents Can Autonomously Hack Websites

LLM Agents Can Autonomously Hack Websites

The paper shows GPT-4 agents with tool use and long context can autonomously exploit real websites, including performing blind SQL injection and schema extraction.

18Agents
Grandmaster-Level Chess Without Search

Grandmaster-Level Chess Without Search

DeepMind shows that a 270M-parameter transformer trained purely with supervised learning on Stockfish-generated data reaches grandmaster-level chess without any search at inference time.

19Training
AnyTool

AnyTool

AnyTool is a training-free LLM agent that scales tool-use to 16K+ Rapid APIs through a hierarchical retriever and a self-reflective solver.

20Agents
Phase Transition in Dot-Product Attention

Phase Transition in Dot-Product Attention

A theoretical paper that analyzes a solvable low-rank tied-QK attention model and uncovers a data-driven phase transition between positional and semantic attention regimes.

21Architecture
Indirect Reasoning with LLMs (DIR)

Indirect Reasoning with LLMs (DIR)

Direct-Indirect Reasoning augments standard CoT with contrapositive and proof-by-contradiction templates, giving LLMs an explicit way to attack problems they can't solve forward.

22Reasoning
ALOHA 2

ALOHA 2

ALOHA 2 is a refreshed low-cost bimanual teleoperation platform from Stanford/DeepMind, designed for large-scale robot-learning data collection.

23Robotics
More Agents Is All You Need

More Agents Is All You Need

The paper shows that simply running more independent LLM agents and voting produces reliable scaling gains across tasks, without any method changes.

24Agents
Self-Discover

Self-Discover

Google's Self-Discover lets LLMs compose their own task-specific reasoning strategies from a small library of atomic reasoning modules, at dramatically lower inference cost than self-consistency.

25Reasoning
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026