Large World Model (LWM)
Free while signed in. Answers cite the passages they came from.

UC Berkeley's LWM is an open 7B multimodal model trained on long videos and books that handles context windows up to 1M tokens via RingAttention.
RingAttention training: Uses blockwise RingAttention to split long sequences across devices in a ring, enabling million-token training at feasible memory cost.
Progressive context expansion: Training context is extended from 4K up to 1M tokens in stages, with masked sequence packing to mix sequence lengths and loss weighting to maintain short-context quality.
Model-generated long-context QA: Uses an LLM to synthesize QA pairs over long inputs for instruction tuning, giving the model targeted practice at answering from deep inside its context.
Open release: Full 7B model family, code, and curated datasets are open-sourced, setting new benchmarks on difficult long-context retrieval and long-video understanding.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack