🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Safety

LIMA

First page
LIMA
Paper summary

Meta's 65B LLaMA fine-tuned on just 1,000 curated examples - showing alignment needs less data than believed.

Ask this paper

Key points
01

1,000-example SFT: Achieves strong alignment with only 1,000 carefully curated prompt-response pairs, no RLHF needed.

02

"Superficial Alignment Hypothesis": Proposes that a model's knowledge is learned in pretraining and alignment mostly teaches response style.

03

GPT-4 competitive: Generates responses preferred over or equivalent to GPT-4 in 43% of cases, and much higher versus Bard.

04

Data-quality over quantity: Became a foundational reference for the "quality over quantity" SFT paradigm that dominated later alignment work.

Every Monday
Get next week’s papers.
Subscribe on Substack