🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
614 papers · ReasoningClear filters →
Gemini 1.0

Gemini 1.0

Google launches Gemini 1.0, a multimodal family natively designed to reason across text, images, video, audio, and code from the ground up.

529Multimodal
LLMs on Graphs

LLMs on Graphs

A comprehensive overview of the many ways LLMs can be applied to graph-structured data and when each pattern is useful.

530Reasoning
Chain of Code

Chain of Code

DeepMind's Chain of Code extends CoT by encouraging LMs to write pseudocode that mixes real code with LM-simulated sub-routines.

531Reasoning
Open-Source LLMs vs. ChatGPT

Open-Source LLMs vs. ChatGPT

A survey cataloguing tasks where open-source LLMs claim to be on par with or better than ChatGPT.

532Evaluation
Medprompt

Medprompt

Microsoft researchers show that careful prompt engineering can push general-purpose GPT-4 to state-of-the-art on medical benchmarks, no domain fine-tuning required.

533Evaluation
System 2 Attention (S2A)

System 2 Attention (S2A)

Meta's S2A uses the LLM's own reasoning to decide what context actually matters, regenerating a clean prompt before the final response step.

534Reasoning
Teaching Small LMs to Reason

Teaching Small LMs to Reason

An approach that teaches smaller language models to explicitly select among reasoning techniques for each problem.

535Reasoning
GPQA

GPQA

A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.

536Reasoning
Hitchhiker's Guide From CoT to Agents

Hitchhiker's Guide From CoT to Agents

A survey mapping the conceptual evolution from chain-of-thought reasoning to modern language-agent frameworks.

537Agents
GAIA

GAIA

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

538Agents
MedAgents

MedAgents

A collaborative multi-round framework for medical reasoning that uses role-playing LLM agents to improve accuracy and reasoning depth.

539Reasoning
LLMs for Scientific Discovery

LLMs for Scientific Discovery

A broad evaluation of GPT-4 across scientific disciplines including drug discovery, biology, and computational chemistry.

540Evaluation
Contrastive Chain-of-Thought

Contrastive Chain-of-Thought

Proposes contrastive CoT prompting where models see both valid *and* invalid reasoning demonstrations to reduce reasoning errors.

541Reasoning
Survey on Language Models for Code

Survey on Language Models for Code

A comprehensive survey of LLMs for code covering 50+ models, 30+ evaluation tasks, and 500 related works.

542Evaluation
LLMs Can Deceive Users (Trading Agent)

LLMs Can Deceive Users (Trading Agent)

Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

543Agents
Rephrase and Respond (RaR)

Rephrase and Respond (RaR)

An effective prompting method where the LLM rephrases and expands the user's question before answering it.

544Reasoning
On the Road with GPT-4V

On the Road with GPT-4V

An exhaustive evaluation of GPT-4V applied to autonomous driving scenarios.

545Evaluation
Evaluating LLMs Survey

Evaluating LLMs Survey

A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.

546Evaluation
EmotionPrompt

EmotionPrompt

Microsoft researchers show that appending emotional stimuli to prompts reliably improves LLM performance across 45 tasks.

547Reasoning
LLMs Meet New Knowledge

LLMs Meet New Knowledge

A benchmark that evaluates how well LLMs handle new knowledge beyond their training cutoff.

548Evaluation
Branch-Solve-Merge (BSM)

Branch-Solve-Merge (BSM)

BSM decomposes LLM tasks into parallel sub-tasks via three LLM-programmed modules: branch, solve, and merge.

549Agents
Llemma

Llemma

Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

550Data
LLMs for Software Engineering

LLMs for Software Engineering

A comprehensive survey of LLMs for software engineering covering models, tasks, evaluation, and open challenges.

551Code
Self-RAG

Self-RAG

Self-RAG trains an LM to adaptively retrieve, generate, and self-critique using special reflection tokens.

552Retrieval
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026