🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture · Data

Attention as an RNN

First page
Attention as an RNN
Paper summary

presents a new attention mechanism that can be trained in parallel (like Transformers) and be updated efficiently with new tokens requiring constant memory usage for inferences (like RNNs); the attention formulation is based on the parallel prefix scan algorithm which enables efficient computation of attention’s many-to-many RNN output; achieves comparable performance to Transformers on 38 datasets while being more time and memory-efficient.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack