LLM2LLM
Free while signed in. Answers cite the passages they came from.

LLM2LLM is an iterative data augmentation scheme where a strong teacher LLM generates new training examples targeted at the specific mistakes a student model makes during fine-tuning.
Error-targeted augmentation: The student trains on a seed dataset, its wrong answers are isolated, and the teacher LLM synthesizes new examples around those failure modes rather than expanding the data uniformly.
Iterative loop: Each round the student is re-trained, re-evaluated, and re-augmented, amplifying signal on hard examples over successive cycles.
Large gains on small data: Using Llama-2-7B, LLM2LLM improves performance by up to 52.6% on TREC and 24.2% on GSM8K over plain fine-tuning baselines in low-data regimes.
Five-dataset evaluation: Strong results are shown on GSM8K, CaseHOLD, SNIPS, TREC, and SST-2, covering math, legal, intent, question, and sentiment classification respectively.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack