The False Promise of Imitating Proprietary LLMs
First page

Paper summary
Berkeley's critical analysis of open-source imitation of proprietary LLMs.
Ask this paper
01
Imitation limits: Shows that fine-tuning small open models on GPT-4 outputs creates a stylistic illusion without meaningfully improving factual capabilities.
02
Stylistic mimicry: Imitation models learn to sound like GPT-4 but retain the base model's underlying capability ceiling.
03
Base model leverage: Argues the higher-leverage action for open-source is building better base models, not imitating proprietary outputs.
04
Field-redirecting: Shifted open-source research focus from distillation toward better pretraining data and scale, preparing the ground for strong foundation models like Llama 2.