🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
AnomalyGPT

AnomalyGPT

Applies large vision-language models to industrial anomaly detection with synthetic data augmentation.

505Data
LegalBench

LegalBench

A collaboratively constructed benchmark for measuring legal reasoning in LLMs.

506Evaluation
Platypus

Platypus

Platypus is a family of fine-tuned and merged LLMs that topped the Open LLM Leaderboard in August 2023.

507Training
Model Compression for LLMs Survey

Model Compression for LLMs Survey

A survey of recent model-compression techniques applied specifically to LLMs.

508Efficiency
GEARS

GEARS

Stanford's GEARS predicts cellular responses to genetic perturbation using deep learning + a gene-relationship knowledge graph.

509Evaluation
Shepherd

Shepherd

Meta's Shepherd is a 7B language model specifically tuned to critique model outputs and suggest refinements.

510Evaluation
Political Biases in NLP Models

Political Biases in NLP Models

Develops methods to measure political and media biases in LLMs and their downstream effects.

511Evaluation
AgentBench

AgentBench

Tsinghua's AgentBench is a multidimensional benchmark for LLM-as-Agent reasoning and decision-making across 8 environments.

512Agents
Trustworthy LLMs

Trustworthy LLMs

Presents a comprehensive framework of categories for assessing LLM trustworthiness.

513Safety
ToolLLM

ToolLLM

Tsinghua's ToolLLM enables LLMs to interact with 16,000+ real-world APIs through a comprehensive framework for tool-using LLMs.

514Agents
Self-Check

Self-Check

Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

515Reasoning
Med-PaLM Multimodal

Med-PaLM Multimodal

Introduces a generalist biomedical AI system and a new multimodal biomedical benchmark with 14 tasks.

516Multimodal
L-Eval

L-Eval

A standardized evaluation suite for long-context language models.

517Evaluation
FacTool

FacTool

A task- and domain-agnostic framework for factuality detection of LLM-generated text.

518Evaluation
How is ChatGPT's Behavior Changing Over Time?

How is ChatGPT's Behavior Changing Over Time?

Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

519Evaluation
FLASK

FLASK

Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.

520Evaluation
Claude 2

Claude 2

Anthropic's second-generation LLM with a detailed model card on safety, alignment, and capabilities.

521Safety
LLMs as General Pattern Machines

LLMs as General Pattern Machines

Demonstrates LLMs serve as general sequence modelers without additional training.

522Reasoning
A Survey on Evaluation of LLMs

A Survey on Evaluation of LLMs

A comprehensive overview of evaluation methods covering what, where, and how to evaluate LLMs.

523Evaluation
InterCode

InterCode

A framework treating interactive coding as a reinforcement learning environment.

524Reinforcement Learning
LeanDojo

LeanDojo

An open-source Lean playground consisting of toolkits, data, models, and benchmarks for theorem proving.

525Reasoning
Generative AI for Programming Education

Generative AI for Programming Education

Evaluates GPT-4 and ChatGPT on programming education scenarios versus human tutors.

526Evaluation
Understanding Theory-of-Mind in LLMs with LLMs

Understanding Theory-of-Mind in LLMs with LLMs

A framework for procedurally generating ToM evaluations using LLMs themselves.

527Evaluation
Evaluations with No Labels

Evaluations with No Labels

Self-supervised evaluation of LLMs via sensitivity/invariance to input transformations.

528Evaluation
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026