OLMo

Allen AI releases OLMo, a truly open 7B-parameter LLM shipped with training code, pretraining data, full weights, evaluation tooling, and fine-tuning recipes - an answer to the "open-weights but closed-pipeline" releases dominating the space.
Ask this paper
Full transparency: Alongside the 7B model, the release includes the Dolma pretraining corpus, the exact training code, intermediate checkpoints, and evaluation harnesses - enabling end-to-end reproducibility that is rare among large open models.
Strong generative performance: OLMo 7B is competitive with Llama 2 and MPT at the same parameter count across generative tasks while being more accessible for downstream research and ablation.
Smaller sibling: A 1B-parameter OLMo 1B is released in parallel, aimed at research on small-model scaling laws and on-device experimentation.
Research enablement: Explicitly positioned as a platform for the community to study what goes into LLM training - data mixing, tokenization, training dynamics - rather than treating the model as the artifact.