🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Efficiency · Architecture

Fast Inference of Mixture-of-Experts

Free while signed in. Answers cite the passages they came from.

First page
Fast Inference of Mixture-of-Experts
The curator’s take

Achieves practical Mixtral-8x7B inference on consumer hardware through MoE-aware quantization and offloading.

Key points
01

Split quantization: Applies different quantization schemes to attention layers vs. expert layers, recognizing that experts tolerate aggressive compression while attention doesn't.

02

MoE-specific offloading: Dynamically shuttles experts between GPU and CPU memory based on routing patterns, exploiting the fact that only a subset of experts activates per token.

03

Consumer hardware: Enables running Mixtral-8x7B on a single desktop GPU and even the free tier of Google Colab - previously reserved for multi-GPU server deployments.

04

Access democratization: Practical on-ramp for researchers and hobbyists to experiment with frontier-class MoE models without cloud infrastructure.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack