The False Promise of Imitating Proprietary LLMs
Free while signed in. Answers cite the passages they came from.

Berkeley's critical analysis of open-source imitation of proprietary LLMs.
Imitation limits: Shows that fine-tuning small open models on GPT-4 outputs creates a stylistic illusion without meaningfully improving factual capabilities.
Stylistic mimicry: Imitation models learn to sound like GPT-4 but retain the base model's underlying capability ceiling.
Base model leverage: Argues the higher-leverage action for open-source is building better base models, not imitating proprietary outputs.
Field-redirecting: Shifted open-source research focus from distillation toward better pretraining data and scale, preparing the ground for strong foundation models like Llama 2.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack