🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,762
Papers
176
Weekly issues
2023
Since
862 papers · EvaluationClear filters →
The Power of Noise: Redefining Retrieval in RAG

The Power of Noise: Redefining Retrieval in RAG

A study stress-testing the retriever component of RAG systems with surprising results about what actually helps generation.

673Retrieval
Hallucination in LVLMs

Hallucination in LVLMs

A survey specifically scoped to hallucination in Large Vision-Language Models, a phenomenon that differs substantially from text-only LLM hallucination.

674Safety
SliceGPT

SliceGPT

Microsoft's SliceGPT is a post-training LLM compression technique that literally slices rows and columns out of weight matrices while preserving zero-shot quality.

675Efficiency
Depth Anything

Depth Anything

A robust monocular depth estimator designed to handle "any image under any circumstance" by scaling self-training on unlabeled data rather than hunting for bigger labeled sets.

676Training
Knowledge Fusion of LLMs (FuseLLM)

Knowledge Fusion of LLMs (FuseLLM)

FuseLLM proposes fusing the capabilities of multiple existing LLMs into a single target model by distilling their output distributions rather than retraining from scratch.

677Training
MambaByte

MambaByte

MambaByte adapts the Mamba state-space architecture to learn directly from raw bytes, bypassing tokenization and all its well-known failure modes.

678Architecture
Resource-efficient LLMs & Multimodal Foundation Models

Resource-efficient LLMs & Multimodal Foundation Models

A wide-ranging survey of efficiency techniques for LLMs and multimodal foundation models, spanning architecture, algorithms, and system design.

679Multimodal
Red Teaming Visual Language Models

Red Teaming Visual Language Models

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.

680Evaluation
Lumiere

Lumiere

Google's Lumiere is a space-time diffusion model for text-to-video that generates the entire video duration in a single forward pass rather than cascading short clips.

681Multimodal
AgentBoard

AgentBoard

AgentBoard is a benchmark and open-source evaluation framework for analytically evaluating LLM agents beyond the usual pass/fail metrics.

682Evaluation
Self-Rewarding Language Models

Self-Rewarding Language Models

Meta shows that an LLM can act as both actor and judge in its own alignment loop, generating training data without any external reward model.

683Reinforcement Learning
Overview of LLMs for Evaluation

Overview of LLMs for Evaluation

A thorough survey of LLM-as-a-Judge and LLM-based evaluation methodologies, mapping strengths, limitations, and open problems.

684Evaluation
Easy-to-Hard Generalization

Easy-to-Hard Generalization

UNC researchers show that LLMs often generalize well from easy training data to hard evaluation data, with implications for scalable oversight.

685Evaluation
Blending Is All You Need

Blending Is All You Need

Small chat models (6B/13B) blended together can rival ChatGPT-class systems, without any new training.

686Evaluation
MagicVideo-V2

MagicVideo-V2

ByteDance's MagicVideo-V2 is an end-to-end text-to-video pipeline that stitches together four specialized modules into a high-fidelity generation system.

687Multimodal
TrustLLM (Trustworthiness in LLMs)

TrustLLM (Trustworthiness in LLMs)

A 100+ page study that defines a principled framework for trustworthy LLMs and benchmarks 16 mainstream models across it.

688Evaluation
Chain-of-Table

Chain-of-Table

Google's Chain-of-Table prompts LLMs to iteratively transform a complex table step-by-step to answer questions reliably, extending CoT reasoning to tabular data.

689Reasoning
RAISE

RAISE

RAISE is an advanced agent architecture that adds a dual-memory system on top of a ReAct-style backbone to better support long-running conversational agents.

690Memory
Quantifying Prompt-Format Sensitivity

Quantifying Prompt-Format Sensitivity

CMU researchers show that LLM few-shot performance is shockingly sensitive to superficial prompt-formatting choices.

691Evaluation
Adversarial Machine Learning (NIST)

Adversarial Machine Learning (NIST)

NIST's official taxonomy of adversarial machine learning, intended to standardize terminology for policy and practice.

692Evaluation
Mitigating Hallucination in LLMs

Mitigating Hallucination in LLMs

A survey cataloging 32 hallucination-mitigation techniques and organizing them into a practical taxonomy.

693Safety
LLaMA Pro

LLaMA Pro

LLaMA Pro introduces block expansion as a recipe for adding new knowledge to a pretrained LLM without catastrophic forgetting.

694Training
SeeAct (GPT-4V as Generalist Web Agent)

SeeAct (GPT-4V as Generalist Web Agent)

OSU researchers adapt GPT-4V into SeeAct, a generalist agent that operates live websites using vision + language planning.

695Agents
DocLLM

DocLLM

JPMorgan's DocLLM is a lightweight extension to LLMs for visual-document understanding that uses bounding-box spatial information rather than image pixels.

696Training
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026