LIMA
First page

Paper summary
Meta's 65B LLaMA fine-tuned on just 1,000 curated examples - showing alignment needs less data than believed.
Ask this paper
01
1,000-example SFT: Achieves strong alignment with only 1,000 carefully curated prompt-response pairs, no RLHF needed.
02
"Superficial Alignment Hypothesis": Proposes that a model's knowledge is learned in pretraining and alignment mostly teaches response style.
03
GPT-4 competitive: Generates responses preferred over or equivalent to GPT-4 in 43% of cases, and much higher versus Bard.
04
Data-quality over quantity: Became a foundational reference for the "quality over quantity" SFT paradigm that dominated later alignment work.