🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
Prompt2Model

Prompt2Model

CMU's Prompt2Model automates the path from a natural-language task description to a deployable small special-purpose model.

02Agents
LegalBench

LegalBench

A collaboratively constructed benchmark for measuring legal reasoning in LLMs.

03Evaluation
Language to Rewards for Robotic Skill Synthesis

Language to Rewards for Robotic Skill Synthesis

Google's Language-to-Rewards uses LLMs to define reward parameters for robotic RL.

04Robotics
Humpback (Self-Alignment with Instruction Backtranslation)

Humpback (Self-Alignment with Instruction Backtranslation)

Meta's Humpback automatically generates instruction-tuning data by back-translating web text into plausible instructions.

05Data
Platypus

Platypus

Platypus is a family of fine-tuned and merged LLMs that topped the Open LLM Leaderboard in August 2023.

06Training
Model Compression for LLMs Survey

Model Compression for LLMs Survey

A survey of recent model-compression techniques applied specifically to LLMs.

07Efficiency
GEARS

GEARS

Stanford's GEARS predicts cellular responses to genetic perturbation using deep learning + a gene-relationship knowledge graph.

08Evaluation
Shepherd

Shepherd

Meta's Shepherd is a 7B language model specifically tuned to critique model outputs and suggest refinements.

09Evaluation
GPT-4 Code Interpreter for Math

GPT-4 Code Interpreter for Math

A zero-shot prompting technique for GPT-4 Code Interpreter that dramatically boosts math-reasoning accuracy via code self-verification.

10Reasoning
Teach LLMs to Personalize

Teach LLMs to Personalize

A multitask-learning approach for personalized text generation without relying on predefined user attributes.

11Training
OctoPack

OctoPack

Hugging Face releases OctoPack, a 4TB dataset of Git commits across 350 programming languages for instruction-tuning code LLMs.

12Data
Outlines (Efficient Guided Generation)

Outlines (Efficient Guided Generation)

A library for guided LLM text generation that enforces structural constraints with minimal overhead.

13Efficiency
Bayesian Flow Networks (BFN)

Bayesian Flow Networks (BFN)

Introduces a new class of generative models that combine Bayesian inference with deep learning.

14Architecture
D-Bot (LLMs as Database Administrators)

D-Bot (LLMs as Database Administrators)

Introduces D-Bot, an LLM-based framework that continuously acquires database-administration knowledge from textual sources.

15Agents
Political Biases in NLP Models

Political Biases in NLP Models

Develops methods to measure political and media biases in LLMs and their downstream effects.

16Evaluation
AgentBench

AgentBench

Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

17Agents
Studying LLM Generalization with Influence Functions

Studying LLM Generalization with Influence Functions

Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.

18Safety
NeuroImagen

NeuroImagen

Reconstructs visual stimuli images from EEG signals using latent diffusion, opening new windows into visually-evoked brain activity.

19Multimodal
SynJax

SynJax

DeepMind's SynJax is a JAX-based library for efficient vectorized inference in structured distributions.

20Efficiency
Synthetic Data Reduces Sycophancy

Synthetic Data Reduces Sycophancy

Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.

21Data
PUG (Photorealistic Unreal Graphics)

PUG (Photorealistic Unreal Graphics)

Meta's PUG uses Unreal Engine to generate photorealistic, semantically controllable synthetic datasets for vision research.

22Data
LLMs for HVAC Control

LLMs for HVAC Control

Microsoft applies LLMs to industrial control tasks (HVAC for buildings), comparing against RL baselines.

23Agents
Trustworthy LLMs

Trustworthy LLMs

Presents a comprehensive framework of categories for assessing LLM trustworthiness.

24Safety
Open Problems and Limitations of RLHF

Open Problems and Limitations of RLHF

A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

25Reinforcement Learning
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026