🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture · Reasoning · Memory

Think Harder or Know More

First page
Think Harder or Know More
Paper summary

This paper investigates transformer models featuring both adaptive per-layer looping, where each block learns to iterate its hidden state via a learned halting mechanism, and gated memory banks that provide additional learned storage. The key finding is that looping primarily benefits mathematical reasoning while memory banks help recover performance on commonsense tasks. Combining both mechanisms yields a model that outperforms an iso-FLOP baseline with three times the number of layers on math benchmarks. Analysis of model internals reveals layer specialization: early layers loop minimally and access memory sparingly, while later layers do both more heavily.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack