ReFT: Representation Finetuning for LMs

Stanford's ReFT freezes the base model and instead learns small interventions on hidden representations at selected layers, offering a more parameter-efficient alternative to LoRA-style PEFT.
Ask this paper
Representations as targets: Instead of updating weights, ReFT trains lightweight linear interventions that modify a rank-limited subspace of the hidden state at specified layers and positions.
LoReFT variant: Low-rank LoReFT is the headline method and is drop-in compatible with the PEFT ecosystem, acting as a direct LoRA replacement.
15-65x fewer parameters than LoRA: Across commonsense reasoning, arithmetic, instruction tuning, and GLUE, LoReFT matches or beats LoRA while using 15-65x fewer trainable parameters.
Interpretability-informed: The method is motivated by interpretability results showing that semantic content is encoded in compact subspaces of the hidden state, and it exploits that structure directly.