MoE-Mamba
Free while signed in. Answers cite the passages they came from.

MoE-Mamba combines state-space models (Mamba) with Mixture-of-Experts to scale LLMs more efficiently than either Mamba or Transformer-MoE alone.
SSM + MoE hybrid: Interleaves Mamba blocks with sparse MoE blocks, combining Mamba's linear-time sequence modeling with MoE's parameter scaling.
2.2x training speedup: Reaches the same loss as plain Mamba in 2.2x fewer training steps at equal parameter count, while preserving Mamba's inference-time advantages over transformers.
Beats Transformer-MoE: Outperforms Transformer-MoE baselines at comparable compute, suggesting that the SSM backbone is a better partner for sparse experts than attention for many long-context workloads.
Scaling signal: Serves as an early proof-of-concept that sparse-expert scaling is architecture-agnostic and will be a key lever for non-transformer foundation models going forward.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack