🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
Is Cosine-Similarity Really About Similarity?

Is Cosine-Similarity Really About Similarity?

This paper argues that cosine similarity between learned embeddings does not always measure semantic similarity, and gives analytical examples where it produces arbitrary or non-unique values.

457Evaluation
Claude 3

Claude 3

Anthropic releases the Claude 3 family (Haiku, Sonnet, Opus), with Opus leapfrogging GPT-4 on many standard benchmarks and bringing frontier multimodal capability plus a much larger context window.

458Evaluation
Robust Evaluation of Reasoning

Robust Evaluation of Reasoning

The paper introduces functional benchmarks that parameterize reasoning problems so the same structural question can be re-instantiated with fresh surface forms, then uses them to expose a large "reasoning gap" in frontier LLMs.

459Reasoning
Can LLMs Reason and Plan?

Can LLMs Reason and Plan?

Kambhampati's position paper argues that what looks like reasoning and planning in LLMs is better understood as "universal approximate retrieval" powered by web-scale training.

460Reasoning
Design2Code

Design2Code

Design2Code tackles the front-end engineering problem of turning a visual design into working HTML/CSS and gives the community both a benchmark and strong MLLM baselines.

461Evaluation
Mistral Large

Mistral Large

Mistral AI releases Mistral Large, its flagship closed-weight LLM positioned as the second-ranked API-accessible model behind GPT-4 at launch.

462Agents
LLMs on Tabular Data: A Survey

LLMs on Tabular Data: A Survey

A survey that maps how LLMs are being applied to tabular data tasks - a domain historically dominated by gradient-boosted trees and specialized architectures.

463Evaluation
Gemini 1.5

Gemini 1.5

Google DeepMind's Gemini 1.5 is a multimodal MoE LLM that scales context to 1M tokens (10M in research settings) while matching or surpassing Gemini 1.0 Ultra on standard benchmarks.

464Memory
ChemLLM

ChemLLM

ChemLLM is a chemistry-specialized LLM with a matched dataset (ChemData) and benchmark (ChemBench) for evaluating chemistry-specific capability.

465Evaluation
Survey of LLMs

Survey of LLMs

A survey that maps the landscape of the three dominant LLM families - GPT, Llama, and PaLM - and the shared toolbox used to build and augment them.

466Evaluation
LLMs for Table Processing: A Survey

LLMs for Table Processing: A Survey

A survey covering how LLMs and VLMs are used across the full spectrum of table-processing tasks, from classic TableQA to spreadsheet manipulation.

467Evaluation
Red Teaming Visual Language Models

Red Teaming Visual Language Models

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.

468Evaluation
AgentBoard

AgentBoard

AgentBoard is a benchmark and open-source evaluation framework for analytically evaluating LLM agents beyond the usual pass/fail metrics.

469Evaluation
Self-Rewarding Language Models

Self-Rewarding Language Models

Meta shows that an LLM can act as both actor and judge in its own alignment loop, generating training data without any external reward model.

470Reinforcement Learning
Overview of LLMs for Evaluation

Overview of LLMs for Evaluation

A thorough survey of LLM-as-a-Judge and LLM-based evaluation methodologies, mapping strengths, limitations, and open problems.

471Evaluation
Easy-to-Hard Generalization

Easy-to-Hard Generalization

UNC researchers show that LLMs often generalize well from easy training data to hard evaluation data, with implications for scalable oversight.

472Evaluation
TrustLLM (Trustworthiness in LLMs)

TrustLLM (Trustworthiness in LLMs)

A 100+ page study that defines a principled framework for trustworthy LLMs and benchmarks 16 mainstream models across it.

473Evaluation
Quantifying Prompt-Format Sensitivity

Quantifying Prompt-Format Sensitivity

CMU researchers show that LLM few-shot performance is shockingly sensitive to superficial prompt-formatting choices.

474Evaluation
LLaMA Pro

LLaMA Pro

LLaMA Pro introduces block expansion as a recipe for adding new knowledge to a pretrained LLM without catastrophic forgetting.

475Training
CogAgent

CogAgent

Tsinghua's CogAgent is an 18B-parameter visual-language model purpose-built for GUI understanding and navigation, with unusually high input resolution.

476Evaluation
PromptBench

PromptBench

A unified library for comprehensive evaluation and analysis of LLMs that consolidates multiple evaluation concerns under one roof.

477Evaluation
Survey of Reasoning with Foundation Models

Survey of Reasoning with Foundation Models

A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

478Reasoning
Gemini's Language Abilities

Gemini's Language Abilities

CMU's impartial, reproducible evaluation of Gemini Pro against GPT and Mixtral across standard LLM benchmarks.

479Evaluation
Mathematical LLMs Survey

Mathematical LLMs Survey

A survey on the progress of LLMs on mathematical reasoning tasks, covering methods, benchmarks, and open problems.

480Reasoning
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026