🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Architecture

DBRX

Free while signed in. Answers cite the passages they came from.

Paper preview
DBRX
The curator’s take

Databricks releases DBRX, a 132B-total / 36B-active open Mixture-of-Experts LLM that beats established open models on MMLU, HumanEval, and GSM8K while delivering 2x faster inference than LLaMA2-70B.

Key points
01

Fine-grained MoE: 16 experts with top-4 selection per token gives 65x more expert combinations than typical MoE configurations, improving capacity without increasing active compute.

02

Pretraining: Trained on 12T carefully curated text-and-code tokens with a 32K context window; the base model is shipped alongside DBRX Instruct.

03

Benchmark wins: DBRX Instruct reaches 73.7% MMLU and 70.1% HumanEval vs 69.8% / 31.0% for LLaMA2-70B; it also edges out Mixtral Instruct on Open LLM Leaderboard composites (74.5% vs 72.7%) and rivals Grok-1 at 2.4x fewer parameters.

04

Practical inference: Serves up to 150 tok/s/user on Databricks Model Serving, roughly 2x faster than LLaMA2-70B, and even outperforms CodeLLaMa-70 Instruct on code tasks despite being general-purpose.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack