🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
258 papers · SafetyClear filters →
Extracting Interpretable Features from Claude 3 Sonnet

Extracting Interpretable Features from Claude 3 Sonnet

presents an effective method to extract millions of abstract features from an LLM that represent specific concepts; these concepts could represent people, places, programming abstractions, emotion, and more; reports that some of the discovered features are directly related to the safety aspects of the model; finds features directly related to security vulnerabilities and backdoors in code, bias, deception, sycophancy; and dangerous/criminal content, and more; these features are also used to intuititively steer the model’s output.

193Safety
Risks and Opportunities of Open-Source Generative AI

Risks and Opportunities of Open-Source Generative AI

analyzes the risks and opportunities of open-source generative AI models; argues that the overall benefits of open-source generative AI outweigh its risks.

194Safety
Fine-tuning and Hallucinations

Fine-tuning and Hallucinations

studies the impact of fine-tuning on new knowledge on the hallucination tendencies of LLMs; the setup includes fine-tuning examples that include new knowledge; shows that LLMs struggle to acquire new factual knowledge via fine-tuning; also finds that as new knowledge is learned it increases the model’s tendency to hallucinate.

195Training
Inner Workings of Transformer Language Models

Inner Workings of Transformer Language Models

presents a technical introduction to current techniques used to interpret the inner workings of Transformer-based language models; it provides a detailed overview of the internal mechanisms implemented in these models.

196Safety
Multimodal LLM Hallucinations

Multimodal LLM Hallucinations

provides an overview of the recent advances in identifying, evaluating, and mitigating hallucination in multimodal LLMs; it also provides an overview of causes, evaluation benchmarks, metrics, and other strategies to deal with challenges related to detecting hallucinations.

197Multimodal
Reducing Hallucination in Structured Outputs via RAG

Reducing Hallucination in Structured Outputs via RAG

This paper deploys a compact RAG pipeline - small retriever plus small LM - for an enterprise workflow-generation task and shows it reduces hallucination while improving out-of-domain generalization vs a baseline LLM.

198Retrieval
Best Practices and Lessons on Synthetic Data

Best Practices and Lessons on Synthetic Data

Google DeepMind's survey-style position paper on synthetic data for LLMs. It covers applications, quality-assurance principles, and the open challenges of factuality, fidelity, bias, and privacy.

199Data
Overview of Multilingual LLMs

Overview of Multilingual LLMs

A first-of-its-kind survey on multilingual LLMs, organized by multilingual alignment principles rather than model-family hierarchy. The authors propose a unified taxonomy and collect open resources to accelerate future research.

200Safety
Many-shot Jailbreaking

Many-shot Jailbreaking

Anthropic shows that long-context windows enable a new attack where hundreds of fake user/assistant dialogues are packed into a single prompt, coaxing frontier LLMs to answer the final harmful question despite safety training.

201Safety
Logits of API-Protected LLMs Leak Proprietary Information

Logits of API-Protected LLMs Leak Proprietary Information

The paper shows that the softmax bottleneck in modern LLMs means even logit-level APIs leak enough information to reconstruct hidden architectural details.

202Safety
Stealing Part of a Production Language Model

Stealing Part of a Production Language Model

The paper demonstrates the first practical attack that extracts the embedding-projection layer of production LLMs through their ordinary logit APIs.

203Safety
On the Societal Impact of Open Foundation Models

On the Societal Impact of Open Foundation Models

Stanford CRFM's policy paper proposes a rigorous framework for assessing the *marginal* risk of open-weight foundation models relative to closed models and pre-existing technologies.

204Safety
Red Teaming Visual Language Models

Red Teaming Visual Language Models

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.

205Evaluation
Patchscopes

Patchscopes

Patchscopes is a general framework for inspecting and intervening on LLM internals by "patching" hidden representations into a second inference pass.

206Safety
Sleeper Agents

Sleeper Agents

Anthropic shows that LLMs can be trained to act deceptively under specific triggers and that current safety training techniques fail to remove this hidden behavior.

207Safety
TrustLLM (Trustworthiness in LLMs)

TrustLLM (Trustworthiness in LLMs)

A 100+ page study that defines a principled framework for trustworthy LLMs and benchmarks 16 mainstream models across it.

208Evaluation
Persuasive Adversarial Prompts (PAP)

Persuasive Adversarial Prompts (PAP)

Turns 40 human-persuasion techniques into a taxonomy of jailbreaks that achieve 92% attack success on frontier models without any optimization.

209Safety
Adversarial Machine Learning (NIST)

Adversarial Machine Learning (NIST)

NIST's official taxonomy of adversarial machine learning, intended to standardize terminology for policy and practice.

210Safety
Mitigating Hallucination in LLMs

Mitigating Hallucination in LLMs

A survey cataloging 32 hallucination-mitigation techniques and organizing them into a practical taxonomy.

211Safety
Exploiting Novel GPT-4 APIs

Exploiting Novel GPT-4 APIs

A red-team study of three newer GPT-4 API surfaces - fine-tuning, function calling, and knowledge retrieval - that reveals each introduces new attack vectors.

212Training
Fact Recalling in LLMs

Fact Recalling in LLMs

A mechanistic-interpretability study showing that early MLP layers function as a lookup table for factual recall.

213Safety
Adversarial Attacks on GPT-4

Adversarial Attacks on GPT-4

Demonstrates that a trivially simple random-search procedure can jailbreak GPT-4 with high reliability.

214Safety
FunSearch

FunSearch

DeepMind's FunSearch uses LLMs as a mutation operator in an evolutionary loop to discover genuinely new mathematical knowledge.

215Safety
Weak-to-Strong Generalization

Weak-to-Strong Generalization

OpenAI's superalignment team shows that weak supervisors can still elicit capabilities from much stronger models - a first empirical signal for scalable oversight.

216Training
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026