Tuning Language Models by Proxy
Free while signed in. Answers cite the passages they came from.

Proxy-tuning steers a large frozen LLM by *decoding-time* logit arithmetic using a much smaller fine-tuned model as a "proxy".
Logit-difference steering: At inference, the target base model's logits are shifted by the difference between a small fine-tuned model and its small base version, transferring the fine-tune's behavior.
No target-model training: The large target LLM is never touched - useful when weights are closed or fine-tuning is prohibitively expensive.
88% of the gap closed: Applied to Llama 2 70B with a 7B proxy, proxy-tuning closes 88% of the gap between the 70B base and the 70B chat-tuned version.
Versatile: Works for instruction tuning, domain adaptation, and task-specific fine-tuning, offering a lightweight steering knob that scales to models users can't fine-tune directly.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack