🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture

Grok-1

Full-paper indexing in progress
Paper preview
Grok-1
Paper summary

xAI open-sources Grok-1, a 314B-parameter Mixture-of-Experts base model, making it the largest openly released LLM at the time of publication.

Ask this paper

Key points
01

MoE at 314B: The architecture activates 25% of weights (~86B active parameters) per token, positioning Grok-1 between dense 70B models and proprietary trillion-parameter systems.

02

Base model release: xAI releases the raw base model weights and network architecture under Apache 2.0 - no instruction tuning, no RLHF, intentionally leaving fine-tuning to the community.

03

Training data cutoff: Pretraining uses a corpus with an October 2023 knowledge cutoff, so downstream fine-tunes inherit roughly 1-year-old world knowledge.

04

Openness signal: The release significantly raises the bar for "open" large-model weights and pressures other labs, coming shortly before Elon Musk's public criticism of OpenAI's closed approach.

Every Monday
Get next week’s papers.
Subscribe on Substack