🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency

Relaxed Recursive Transformers

First page
Relaxed Recursive Transformers
Paper summary

introduces a novel approach, Relaxed Recursive Transformer, that significantly reduces LLM size through parameter sharing across layers while maintaining performance; the model is initialized from standard pretrained Transformers, but only uses a single block of unique layers that is repeated multiple times in a loop; then it adds flexibility to the layer tying constraint via depth-wise low-rank adaptation (LoRA) modules; shows that the approach has the potential to lead to significant (2-3×) gains in inference throughput.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack