Cracks in the Foundation
Free while signed in. Answers cite the passages they came from.

You might assume architectural variations within the dense transformer paradigm barely move accuracy, and in the short-context setting you would be right. This work shows four minor decisions, normalization, GQA, pretraining context length, and sliding window attention, each made by at least one of the Olmo, Llama, and Qwen dense families, have a compoundingly negative effect on long-context extensibility. Any one alone is minor, but combining three or more drops downstream long-context performance by up to 47%, and none of it is detectable from short-context loss or validation sets, which is precisely how these choices survive into shipped models. Applying context extension early in pretraining exposes the problem cheaply. After over 170,000 GPU hours the authors release OlmPool, 26 comparable 7B models with checkpoints before and after extension, including several architectures that beat the Llama 3 architecture on long-context extensibility.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack