πŸš€NEW LABGetting Started with Claude AgentsStart lab
Memory

Large Memory Models

First page
Large Memory Models
Paper summary

Large Memory Models (LM2) is a transformer architecture augmented with an external memory module to tackle tasks requiring extensive reasoning and long context. Key highlights include:

Ask this paper

Key points
01

Memory-augmented transformer: LM2 adds a dedicated memory repository that the model can read/write via cross-attention, enabling it to store and retrieve information across many reasoning steps. This design addresses the limitations of standard transformers in tasks like multi-hop reasoning and relational argumentation.

02

Superior long-term reasoning: On the BABILong benchmark for long-context reasoning, LM2 dramatically outperformed prior models – 37% better than a recurrent memory transformer and 86% better than a baseline Llama model on average. It excels at multi-hop inference, numeric reasoning, and QA over long documents.

03

No trade-off in generality: Impressively, LM2 maintained strong general performance – e.g. a +5% boost on the MMLU knowledge test over a baseline – indicating the memory module helps complex tasks without hurting normal language understanding.

04

Alignment via memory: These results underscore the importance of explicit memory for aligning AI reasoning with complex tasks. By integrating a large-scale memory, we get models that can better adhere to task objectives over long dialogues or reasoning chains, a step forward for building more aligned and capable AI systems.

Every Monday
Get next week’s papers.
Subscribe on Substack