🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
860 papers · EvaluationClear filters →
PromptBench

PromptBench

A unified library for comprehensive evaluation and analysis of LLMs that consolidates multiple evaluation concerns under one roof.

697Evaluation
Exploiting Novel GPT-4 APIs

Exploiting Novel GPT-4 APIs

A red-team study of three newer GPT-4 API surfaces - fine-tuning, function calling, and knowledge retrieval - that reveals each introduces new attack vectors.

698Training
Survey of Reasoning with Foundation Models

Survey of Reasoning with Foundation Models

A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

699Reasoning
LLaRA

LLaRA

LLaRA adapts a decoder-only LLM for dense retrieval via two tailored pretext tasks that leverage text embeddings from the LLM itself.

700Retrieval
Gemini's Language Abilities

Gemini's Language Abilities

CMU's impartial, reproducible evaluation of Gemini Pro against GPT and Mixtral across standard LLM benchmarks.

701Evaluation
RAG for LLMs

RAG for LLMs

A broad survey of Retrieval-Augmented Generation research, organizing the rapidly growing literature into a coherent map.

702Retrieval
BabyLLM Challenge Findings

BabyLLM Challenge Findings

Reports results from a challenge on sample-efficient pretraining using a developmentally plausible corpus.

703Training
Mathematical LLMs Survey

Mathematical LLMs Survey

A survey on the progress of LLMs on mathematical reasoning tasks, covering methods, benchmarks, and open problems.

704Reasoning
LLM360

LLM360

LLM360 is a framework for fully transparent open-source LLM development, with everything from data to training dynamics released.

705Training
LLMs in Medicine

LLMs in Medicine

A comprehensive survey (300+ papers) of LLMs applied to medicine, from clinical tasks to biomedical research.

706Evaluation
Pearl

Pearl

Meta's Pearl is a production-ready reinforcement learning agent package designed for real-world deployment constraints.

707Agents
Gemini 1.0

Gemini 1.0

Google launches Gemini 1.0, a multimodal family natively designed to reason across text, images, video, audio, and code from the ground up.

708Multimodal
RankZephyr

RankZephyr

RankZephyr is an open-source LLM for listwise zero-shot reranking that bridges the effectiveness gap with GPT-4.

709Evaluation
Open-Source LLMs vs. ChatGPT

Open-Source LLMs vs. ChatGPT

A survey cataloguing tasks where open-source LLMs claim to be on par with or better than ChatGPT.

710Evaluation
MEDITRON-70B

MEDITRON-70B

EPFL's MEDITRON is an open-source family of medical LLMs at 7B and 70B parameters, continually pretrained on curated medical corpora.

711Training
Medprompt

Medprompt

Microsoft researchers show that careful prompt engineering can push general-purpose GPT-4 to state-of-the-art on medical benchmarks, no domain fine-tuning required.

712Evaluation
UniIR

UniIR

UniIR is a unified instruction-guided multimodal retriever that handles eight retrieval tasks across modalities with a single model.

713Multimodal
Advancing Long-Context LLMs

Advancing Long-Context LLMs

A survey of methodologies for improving Transformer long-context capability across pretraining, fine-tuning, and inference stages.

714Memory
Mirasol3B

Mirasol3B

Google's Mirasol3B is a multimodal model that decouples modalities into focused autoregressive components rather than forcing a single fused stream.

715Multimodal
GPQA

GPQA

A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.

716Reasoning
GAIA

GAIA

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

717Agents
MedAgents

MedAgents

A collaborative multi-round framework for medical reasoning that uses role-playing LLM agents to improve accuracy and reasoning depth.

718Reasoning
TÜLU 2

TÜLU 2

Allen AI's TÜLU 2 is a suite of improved open instruction-tuned LLMs and an accompanying study of adaptation best practices.

719Training
Chain-of-Note (CoN)

Chain-of-Note (CoN)

Tencent's Chain-of-Note adds an explicit note-taking step to RAG so the model can evaluate retrieved evidence before answering.

720Retrieval
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026