🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency · Architecture

Fast Inference of Mixture-of-Experts

First page
Fast Inference of Mixture-of-Experts
Paper summary

Achieves practical Mixtral-8x7B inference on consumer hardware through MoE-aware quantization and offloading.

Ask this paper

Key points
01

Split quantization: Applies different quantization schemes to attention layers vs. expert layers, recognizing that experts tolerate aggressive compression while attention doesn't.

02

MoE-specific offloading: Dynamically shuttles experts between GPU and CPU memory based on routing patterns, exploiting the fact that only a subset of experts activates per token.

03

Consumer hardware: Enables running Mixtral-8x7B on a single desktop GPU and even the free tier of Google Colab - previously reserved for multi-GPU server deployments.

04

Access democratization: Practical on-ramp for researchers and hobbyists to experiment with frontier-class MoE models without cloud infrastructure.

Every Monday
Get next week’s papers.
Subscribe on Substack