🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training

Tuning Language Models by Proxy

Free while signed in. Answers cite the passages they came from.

First page
Tuning Language Models by Proxy
The curator’s take

Proxy-tuning steers a large frozen LLM by *decoding-time* logit arithmetic using a much smaller fine-tuned model as a "proxy".

Key points
01

Logit-difference steering: At inference, the target base model's logits are shifted by the difference between a small fine-tuned model and its small base version, transferring the fine-tune's behavior.

02

No target-model training: The large target LLM is never touched - useful when weights are closed or fine-tuning is prohibitively expensive.

03

88% of the gap closed: Applied to Llama 2 70B with a 7B proxy, proxy-tuning closes 88% of the gap between the 70B base and the 70B chat-tuned version.

04

Versatile: Works for instruction tuning, domain adaptation, and task-specific fine-tuning, offering a lightweight steering knob that scales to models users can't fine-tune directly.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack