🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
860 papers · EvaluationClear filters →
Survey on Factuality in LLMs

Survey on Factuality in LLMs

A survey covering evaluation and enhancement techniques for LLM factuality.

745Evaluation
LLMs for Healthcare Survey

LLMs for Healthcare Survey

A comprehensive overview of LLMs applied to the healthcare domain.

746Evaluation
InstructRetro

InstructRetro

NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.

747Training
FireAct (Language Agent Fine-tuning)

FireAct (Language Agent Fine-tuning)

Explores fine-tuning LLMs specifically for language-agent use, demonstrating consistent gains over prompting alone.

748Agents
LLMs Represent Space and Time

LLMs Represent Space and Time

MIT researchers find that LLMs internally encode linear representations of space and time across multiple scales.

749Reasoning
Retrieval Meets Long-Context LLMs

Retrieval Meets Long-Context LLMs

NVIDIA's study comparing RAG and long-context LLMs, with the punchline that the two are complementary rather than substitutes.

750Retrieval
RA-DIT (Retrieval-Augmented Dual Instruction Tuning)

RA-DIT (Retrieval-Augmented Dual Instruction Tuning)

Meta's RA-DIT is a lightweight recipe that retrofits LLMs with retrieval capabilities through dual fine-tuning.

751Retrieval
Analogical Prompting

Analogical Prompting

Google's Analogical Prompting guides LLM reasoning by having the model self-generate relevant exemplars on the fly.

752Reasoning
Effective Long-Context Scaling (Meta)

Effective Long-Context Scaling (Meta)

Meta proposes a 70B long-context LLM that surpasses GPT-3.5-turbo-16k on long-context benchmarks.

753Memory
Graph Neural Prompting (GNP)

Graph Neural Prompting (GNP)

A plug-and-play method that injects knowledge-graph information into frozen pretrained LLMs.

754Training
Boolformer

Boolformer

The first Transformer trained to perform end-to-end symbolic regression of Boolean functions.

755Reasoning
LLaVA-RLHF

LLaVA-RLHF

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

756Reinforcement Learning
LLM Alignment Survey

LLM Alignment Survey

A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.

757Safety
Qwen

Qwen

Alibaba releases the Qwen family of open LLMs with strong tool-use and planning capabilities for language agents.

758Training
Logical Chain-of-Thought (LogiCoT)

Logical Chain-of-Thought (LogiCoT)

A neurosymbolic framework that verifies and revises zero-shot CoT reasoning using symbolic-logic principles.

759Reasoning
Contrastive Decoding for Reasoning

Contrastive Decoding for Reasoning

Shows that contrastive decoding, a simple inference-time technique, substantially improves reasoning in large LLMs.

760Reasoning
Struc-Bench (LLMs for Structured Data)

Struc-Bench (LLMs for Structured Data)

Studies how LLMs handle complex structured-data generation and proposes a structure-aware fine-tuning method.

761Evaluation
LMSYS-Chat-1M

LMSYS-Chat-1M

LMSYS releases a large-scale dataset of 1 million real-world LLM conversations collected from the Vicuna demo and Chatbot Arena.

762Data
OWL (LLMs for IT Operations)

OWL (LLMs for IT Operations)

Proposes OWL, an LLM specialized for IT operations through self-instruct fine-tuning on IT-specific tasks.

763Evaluation
Rewindable Auto-regressive INference (RAIN)

Rewindable Auto-regressive INference (RAIN)

Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

764Safety
Hallucination Survey (Early)

Hallucination Survey (Early)

Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

765Safety
Radiology-Llama 2

Radiology-Llama 2

A Llama 2-based LLM specialized for radiology report generation.

766Training
MAmmoTH

MAmmoTH

An open-source LLM family specialized for general mathematical problem solving.

767Reasoning
Overview of AI Deception

Overview of AI Deception

A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

768Safety
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026