StarCoder 2

BigCode releases StarCoder 2, an open family of code LLMs at 3B, 7B, and 15B parameters trained on The Stack v2, a much larger and cleaner code corpus than the original StarCoder.
Ask this paper
Multi-org collaboration: ServiceNow trains the 3B, Hugging Face the 7B, and NVIDIA the 15B, with a shared data pipeline and recipe.
Training setup: 15B variant is trained on 4T+ tokens across 600+ programming languages with a 16K-token context window, grouped-query attention, and a fill-in-the-middle objective.
Benchmark punch above its weight: StarCoder2-15B matches 33B+ models on many code completion, reasoning, and PAL-augmented math tasks, while the 3B matches the original 15B StarCoder.
The Stack v2: Released alongside the models is a 67.5TB code dataset derived from Software Heritage, significantly improving on The Stack v1 in both scale and provenance.