🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training

Grok-1

Free while signed in. Answers cite the passages they came from.

Paper preview
Grok-1
The curator’s take

xAI open-sources Grok-1, a 314B-parameter Mixture-of-Experts base model, making it the largest openly released LLM at the time of publication.

Key points
01

MoE at 314B: The architecture activates 25% of weights (~86B active parameters) per token, positioning Grok-1 between dense 70B models and proprietary trillion-parameter systems.

02

Base model release: xAI releases the raw base model weights and network architecture under Apache 2.0 - no instruction tuning, no RLHF, intentionally leaving fine-tuning to the community.

03

Training data cutoff: Pretraining uses a corpus with an October 2023 knowledge cutoff, so downstream fine-tunes inherit roughly 1-year-old world knowledge.

04

Openness signal: The release significantly raises the bar for "open" large-model weights and pressures other labs, coming shortly before Elon Musk's public criticism of OpenAI's closed approach.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack