🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture

Were RNNs All We Needed?

First page
Were RNNs All We Needed?
Paper summary

revisits RNNs and shows that by removing the hidden states from input, forget, and update gates RNNs can be efficiently trained in parallel; this is possible because with this change architectures like LSTMs and GRUs no longer require backpropagate through time (BPTT); they introduce minLSTMs and minGRUs that are 175x faster for a 512 sequence length.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack