🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture

Differential Transformer

First page
Differential Transformer
Paper summary

proposes a differential attention mechanism that amplifies attention to the relevant context while canceling noise; Differential Transformer outperforms Transformer when scaling up model size and training tokens; the authors claim that since this architecture gets less "distracted" by irrelevant context, it can do well in applications such as long-context modeling, key information retrieval, hallucination mitigation, in-context learning, and reduction of activation outliers.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack