🚀NEW LABGetting Started with Claude AgentsStart lab
Training

Tuning Language Models by Proxy

First page
Tuning Language Models by Proxy
Paper summary

Proxy-tuning steers a large frozen LLM by *decoding-time* logit arithmetic using a much smaller fine-tuned model as a "proxy".

Ask this paper

Key points
01

Logit-difference steering: At inference, the target base model's logits are shifted by the difference between a small fine-tuned model and its small base version, transferring the fine-tune's behavior.

02

No target-model training: The large target LLM is never touched - useful when weights are closed or fine-tuning is prohibitively expensive.

03

88% of the gap closed: Applied to Llama 2 70B with a 7B proxy, proxy-tuning closes 88% of the gap between the 70B base and the 70B chat-tuned version.

04

Versatile: Works for instruction tuning, domain adaptation, and task-specific fine-tuning, offering a lightweight steering knob that scales to models users can't fine-tune directly.

Every Monday
Get next week’s papers.
Subscribe on Substack