🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Data

LLM2LLM

First page
LLM2LLM
Paper summary

LLM2LLM is an iterative data augmentation scheme where a strong teacher LLM generates new training examples targeted at the specific mistakes a student model makes during fine-tuning.

Ask this paper

Key points
01

Error-targeted augmentation: The student trains on a seed dataset, its wrong answers are isolated, and the teacher LLM synthesizes new examples around those failure modes rather than expanding the data uniformly.

02

Iterative loop: Each round the student is re-trained, re-evaluated, and re-augmented, amplifying signal on hard examples over successive cycles.

03

Large gains on small data: Using Llama-2-7B, LLM2LLM improves performance by up to 52.6% on TREC and 24.2% on GSM8K over plain fine-tuning baselines in low-data regimes.

04

Five-dataset evaluation: Strong results are shown on GSM8K, CaseHOLD, SNIPS, TREC, and SST-2, covering math, legal, intent, question, and sentiment classification respectively.

Every Monday
Get next week’s papers.
Subscribe on Substack