LLM2LLM

LLM2LLM is an iterative data augmentation scheme where a strong teacher LLM generates new training examples targeted at the specific mistakes a student model makes during fine-tuning.
Ask this paper
Error-targeted augmentation: The student trains on a seed dataset, its wrong answers are isolated, and the teacher LLM synthesizes new examples around those failure modes rather than expanding the data uniformly.
Iterative loop: Each round the student is re-trained, re-evaluated, and re-augmented, amplifying signal on hard examples over successive cycles.
Large gains on small data: Using Llama-2-7B, LLM2LLM improves performance by up to 52.6% on TREC and 24.2% on GSM8K over plain fine-tuning baselines in low-data regimes.
Five-dataset evaluation: Strong results are shown on GSM8K, CaseHOLD, SNIPS, TREC, and SST-2, covering math, legal, intent, question, and sentiment classification respectively.