On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models

William L. Tong, Eran Malach, Emmanuel Abbe, Cengiz Pehlevan and colleagues (Apple, Harvard) show that the gating mechanism in state space models drives both their weak in-context retrieval and their better length generalization.
Ask this paper
Theory: Gating leads SSMs to learn an in-weights memorization solution first, delaying or preventing convergence to the correct in-context solution even when capacity is sufficient.
Long context: Gating sets an effective context length; without decay the relevant signal competes with every distractor, so some gating helps generalization to longer sequences.
Experiments: Synthetic multi-token retrieval and logical-rule retrieval confirm the prediction, and changing the gate at initialization speeds convergence in some settings.
Tool calling: Finetuning Mamba on multi-tool calling, adjusting the initial gate after pretraining raises or lowers hallucination, and a weaker gate than pretraining learned can improve accuracy.
Abstract
State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute. Although SSMs exhibit reasonable performance and favorable computational characteristics, they continue to lag behind Transformers on tasks that require in-context learning and precise retrieval, slowing their adoption for large-scale language modeling. In this work, we demonstrate that both the success and failure of SSMs in these domains can be explained by studying the role of the gating mechanism, a prevalent component in modern recurrent networks. Specifically, we show through theory and experiments that this gating mechanism causes SSMs to first learn an in-weights "memorization" solution, while delaying, or even preventing, convergence to a correct in-context learning solution. Importantly, this happens even in cases where there are no fundamental limitations due to the architecture or its memory capacity. On the other hand, we find that gating is often beneficial for improving generalization to long sequence lengths. Our results illuminate the crucial role of the gating mechanism in shaping both the training dynamics and generalization of SSMs, and provide a basis for understanding and improving linear-time models.