AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.

The Larger They Are, the Harder They Fail
Reveals inverse-scaling failures in LLM code generation.

Model Evaluation for Extreme Risks
DeepMind's framework for evaluating models for catastrophic-risk capabilities.

LLM Research Directions
A list of research directions for students entering LLM research.

LLM Explains Neurons in LLMs
OpenAI's automated interpretability pipeline using GPT-4 to explain GPT-2 neurons.

Unfaithful Explanations in Chain-of-Thought Prompting
Demonstrates CoT explanations can misrepresent the true reason for a model's prediction.

Interpretable ML for Science with PySR
An open-source library for practical symbolic regression in the sciences.

Poisoning Language Models During Instruction Tuning
Shows adversaries can poison LLMs via instruction tuning data.

ChemCrow: Augmenting LLMs with Chemistry Tools
An LLM chemistry agent with 13 expert-designed tools.

Instruction Tuning with GPT-4
Uses GPT-4 to generate instruction-following data for LLM fine-tuning.

Eight Things to Know about Large Language Models
Sam Bowman's influential primer on key LLM considerations.

A Survey of Large Language Models
A 50-page comprehensive survey on LLMs.

MACHIAVELLI Benchmark
A benchmark of 134 text-based Choose-Your-Own-Adventure games for measuring ethical trade-offs.

Pythia
EleutherAI's suite for analyzing LLMs across training and scaling.