🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Evaluation

LLaMA Pro

First page
LLaMA Pro
Paper summary

LLaMA Pro introduces block expansion as a recipe for adding new knowledge to a pretrained LLM without catastrophic forgetting.

Ask this paper

Key points
01

Block expansion: Additional identity-initialized transformer blocks are inserted into a frozen base model; only these new blocks are trained on the new corpus while inherited blocks stay frozen.

02

Math + code training: LLaMA Pro-8.3B is initialized from Llama 2-7B and post-pretrained on a mix of math and code data, yielding a domain-specialized yet general-purpose model.

03

Preserves general capability: Because inherited blocks stay frozen, the model retains its original general skills - a stark contrast with continual full fine-tuning which typically degrades general performance.

04

Strong benchmark performance: Matches or beats Llama 2 7B on general benchmarks while significantly outperforming it on math and code tasks, validating the block-expansion recipe.

Every Monday
Get next week’s papers.
Subscribe on Substack