🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency

TensorLLM

First page
TensorLLM
Paper summary

Proposes a framework that performs MHA compression through a multi-head tensorisation process and the Tucker decomposition. Achieves a compression rate of up to ∼ 250x in the MHA weights, without requiring any additional data, training, or fine-tuning.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack