🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023Issue 182 · Sep 28 – Oct 4, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
This week · 10 papersView the full issue →
LegalBench

LegalBench

A collaboratively constructed benchmark for measuring legal reasoning in LLMs.

02Evaluation
Language to Rewards for Robotic Skill Synthesis

Language to Rewards for Robotic Skill Synthesis

Google's Language-to-Rewards uses LLMs to define reward parameters for robotic RL.

03Robotics
Humpback (Self-Alignment with Instruction Backtranslation)

Humpback (Self-Alignment with Instruction Backtranslation)

Meta's Humpback automatically generates instruction-tuning data by back-translating web text into plausible instructions.

04Data
Platypus

Platypus

Platypus is a family of fine-tuned and merged LLMs that topped the Open LLM Leaderboard in August 2023.

05Training
Model Compression for LLMs Survey

Model Compression for LLMs Survey

A survey of recent model-compression techniques applied specifically to LLMs.

06Efficiency
GEARS

GEARS

Stanford's GEARS predicts cellular responses to genetic perturbation using deep learning + a gene-relationship knowledge graph.

07Evaluation
Shepherd

Shepherd

Meta's Shepherd is a 7B language model specifically tuned to critique model outputs and suggest refinements.

08Evaluation
GPT-4 Code Interpreter for Math

GPT-4 Code Interpreter for Math

A zero-shot prompting technique for GPT-4 Code Interpreter that dramatically boosts math-reasoning accuracy via code self-verification.

09Reasoning
Teach LLMs to Personalize

Teach LLMs to Personalize

A multitask-learning approach for personalized text generation without relying on predefined user attributes.

10Training
OctoPack

OctoPack

Hugging Face releases OctoPack, a 4TB dataset of Git commits across 350 programming languages for instruction-tuning code LLMs.

11Data
Outlines (Efficient Guided Generation)

Outlines (Efficient Guided Generation)

A library for guided LLM text generation that enforces structural constraints with minimal overhead.

12Efficiency
Bayesian Flow Networks (BFN)

Bayesian Flow Networks (BFN)

Introduces a new class of generative models that combine Bayesian inference with deep learning.

13Architecture
D-Bot (LLMs as Database Administrators)

D-Bot (LLMs as Database Administrators)

Introduces D-Bot, an LLM-based framework that continuously acquires database-administration knowledge from textual sources.

14Agents
Political Biases in NLP Models

Political Biases in NLP Models

Develops methods to measure political and media biases in LLMs and their downstream effects.

15Evaluation
AgentBench

AgentBench

Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

16Agents
Studying LLM Generalization with Influence Functions

Studying LLM Generalization with Influence Functions

Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.

17Safety
NeuroImagen

NeuroImagen

Reconstructs visual stimuli images from EEG signals using latent diffusion, opening new windows into visually-evoked brain activity.

18Multimodal
SynJax

SynJax

DeepMind's SynJax is a JAX-based library for efficient vectorized inference in structured distributions.

19Efficiency
Synthetic Data Reduces Sycophancy

Synthetic Data Reduces Sycophancy

Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.

20Data
PUG (Photorealistic Unreal Graphics)

PUG (Photorealistic Unreal Graphics)

Meta's PUG uses Unreal Engine to generate photorealistic, semantically controllable synthetic datasets for vision research.

21Data
LLMs for HVAC Control

LLMs for HVAC Control

Microsoft applies LLMs to industrial control tasks (HVAC for buildings), comparing against RL baselines.

22Agents
Trustworthy LLMs

Trustworthy LLMs

Presents a comprehensive framework of categories for assessing LLM trustworthiness.

23Safety
Open Problems and Limitations of RLHF

Open Problems and Limitations of RLHF

A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

24Reinforcement Learning
Med-Flamingo

Med-Flamingo

Stanford's Med-Flamingo is a multimodal medical model supporting in-context learning for few-shot medical visual QA.

25Multimodal
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026