🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
Learning to Filter Context for RAG (FILCO)

Learning to Filter Context for RAG (FILCO)

CMU's FILCO improves RAG by training a dedicated model to filter retrieved contexts before they reach the generator.

02Retrieval
MART (Multi-round Automatic Red-Teaming)

MART (Multi-round Automatic Red-Teaming)

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

03Safety
LLMs Can Deceive Users (Trading Agent)

LLMs Can Deceive Users (Trading Agent)

Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

04Agents
Hallucination in LLMs Survey

Hallucination in LLMs Survey

A comprehensive survey of hallucination in LLMs, covering taxonomy, causes, evaluation, and mitigation.

05Safety
Simplifying Transformer Blocks

Simplifying Transformer Blocks

Researchers show that many components of the standard transformer block can be removed with no loss in training speed or quality.

06Architecture
In-Context Learning Generalization Limits

In-Context Learning Generalization Limits

Investigates whether transformers' in-context learning can generalize beyond the distribution of their pretraining data.

07Training
MusicGen

MusicGen

Meta's MusicGen is a single-stage transformer LLM for music generation that operates over compressed discrete audio tokens.

08Multimodal
AltUp (Alternating Updates)

AltUp (Alternating Updates)

Google's AltUp lets transformers benefit from wider representations without paying the full compute cost at every layer.

09Architecture
Rephrase and Respond (RaR)

Rephrase and Respond (RaR)

An effective prompting method where the LLM rephrases and expands the user's question before answering it.

10Reasoning
On the Road with GPT-4V

On the Road with GPT-4V

An exhaustive evaluation of GPT-4V applied to autonomous driving scenarios.

11Evaluation
GPT4All Technical Report

GPT4All Technical Report

The GPT4All technical report documents the model family and the open ecosystem built around democratizing local LLMs.

12Data
S-LoRA

S-LoRA

S-LoRA enables serving thousands of LoRA adapters concurrently on a single GPU through memory-paging and custom CUDA kernels.

13Memory
FreshLLMs (FreshQA)

FreshLLMs (FreshQA)

Introduces FreshQA, a dynamic benchmark designed to stress-test LLMs on time-sensitive knowledge.

14Evaluation
MetNet-3

MetNet-3

Google's MetNet-3 is a state-of-the-art neural weather model extending lead time and variable coverage well beyond prior observation-based models.

15Architecture
Evaluating LLMs Survey

Evaluating LLMs Survey

A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.

16Evaluation
Battle of the Backbones

Battle of the Backbones

A large-scale benchmarking framework that compares vision backbones across a diverse suite of computer vision tasks.

17Architecture
ChipNeMo (LLMs for Chip Design)

ChipNeMo (LLMs for Chip Design)

NVIDIA's ChipNeMo applies domain-adapted LLMs to industrial chip design workflows.

18Training
YaRN (Efficient Context Extension)

YaRN (Efficient Context Extension)

YaRN is a compute-efficient method for extending the context window of LLMs well beyond their pretrained length.

19Training
Open DAC 2023

Open DAC 2023

Meta releases a large DFT dataset for training ML models that predict sorbent-adsorbate interactions in Direct Air Capture (DAC).

20Data
Symmetry in Machine Learning

Symmetry in Machine Learning

A methodological framework for enforcing, discovering, and promoting symmetry in machine learning models.

21Architecture
Next-Generation AlphaFold

Next-Generation AlphaFold

DeepMind previews the next AlphaFold with dramatically expanded scope of biomolecular complexes.

22Architecture
EmotionPrompt

EmotionPrompt

Microsoft researchers show that appending emotional stimuli to prompts reliably improves LLM performance across 45 tasks.

23Reasoning
FP8-LM

FP8-LM

Microsoft's FP8-LM demonstrates that most LLM training variables - gradients, optimizer states - can use FP8 without sacrificing accuracy.

24Efficiency
Zephyr

Zephyr

Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

25Reinforcement Learning
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026