Fast Inference of Mixture-of-Experts
Free while signed in. Answers cite the passages they came from.

Achieves practical Mixtral-8x7B inference on consumer hardware through MoE-aware quantization and offloading.
Split quantization: Applies different quantization schemes to attention layers vs. expert layers, recognizing that experts tolerate aggressive compression while attention doesn't.
MoE-specific offloading: Dynamically shuttles experts between GPU and CPU memory based on routing patterns, exploiting the fact that only a subset of experts activates per token.
Consumer hardware: Enables running Mixtral-8x7B on a single desktop GPU and even the free tier of Google Colab - previously reserved for multi-GPU server deployments.
Access democratization: Practical on-ramp for researchers and hobbyists to experiment with frontier-class MoE models without cloud infrastructure.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack