🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
520 papers · 2024Clear filters →
TrustLLM (Trustworthiness in LLMs)

TrustLLM (Trustworthiness in LLMs)

A 100+ page study that defines a principled framework for trustworthy LLMs and benchmarks 16 mainstream models across it.

505Evaluation
Chain-of-Table

Chain-of-Table

Google's Chain-of-Table prompts LLMs to iteratively transform a complex table step-by-step to answer questions reliably, extending CoT reasoning to tabular data.

506Reasoning
Persuasive Adversarial Prompts (PAP)

Persuasive Adversarial Prompts (PAP)

Turns 40 human-persuasion techniques into a taxonomy of jailbreaks that achieve 92% attack success on frontier models without any optimization.

507Safety
RAISE

RAISE

RAISE is an advanced agent architecture that adds a dual-memory system on top of a ReAct-style backbone to better support long-running conversational agents.

508Memory
Quantifying Prompt-Format Sensitivity

Quantifying Prompt-Format Sensitivity

CMU researchers show that LLM few-shot performance is shockingly sensitive to superficial prompt-formatting choices.

509Evaluation
Adversarial Machine Learning (NIST)

Adversarial Machine Learning (NIST)

NIST's official taxonomy of adversarial machine learning, intended to standardize terminology for policy and practice.

510Evaluation
Mobile ALOHA

Mobile ALOHA

Stanford's Mobile ALOHA is a low-cost bimanual mobile-manipulation platform that learns dexterous household tasks via whole-body teleoperation and behavior cloning.

511Robotics
Mitigating Hallucination in LLMs

Mitigating Hallucination in LLMs

A survey cataloging 32 hallucination-mitigation techniques and organizing them into a practical taxonomy.

512Safety
Self-Play Fine-Tuning (SPIN)

Self-Play Fine-Tuning (SPIN)

SPIN shows that a supervised fine-tuned LLM can keep improving via self-play alone, without any additional human annotations.

513Training
LLaMA Pro

LLaMA Pro

LLaMA Pro introduces block expansion as a recipe for adding new knowledge to a pretrained LLM without catastrophic forgetting.

514Training
LLM Augmented LLMs (CALM)

LLM Augmented LLMs (CALM)

Google's CALM composes a large anchor LLM with smaller specialist models via learned cross-attention, unlocking new capabilities without retraining either model.

515Code
Fast Inference of Mixture-of-Experts

Fast Inference of Mixture-of-Experts

Achieves practical Mixtral-8x7B inference on consumer hardware through MoE-aware quantization and offloading.

516Efficiency
SeeAct (GPT-4V as Generalist Web Agent)

SeeAct (GPT-4V as Generalist Web Agent)

OSU researchers adapt GPT-4V into SeeAct, a generalist agent that operates live websites using vision + language planning.

517Agents
DocLLM

DocLLM

JPMorgan's DocLLM is a lightweight extension to LLMs for visual-document understanding that uses bounding-box spatial information rather than image pixels.

518Training
How Code Empowers LLMs

How Code Empowers LLMs

A survey on why training LLMs with code data produces capabilities well beyond coding itself.

519Agents
Instruct-Imagen

Instruct-Imagen

Google's Instruct-Imagen is a multimodal instruction-tuned image generation model that generalizes across heterogeneous generation tasks, including unseen ones.

520Multimodal
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026