Mixture of Experts: Sparse by Design
Routers, experts, and load balancing — how MoE models activate a fraction of their parameters per token, and what that costs in memory.
Content last verified 2026-09.
Lessons
Sources
- Shazeer et al. (2017) — Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Fedus, Zoph & Shazeer (2021) — Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Jiang et al. (2024) — Mixtral of Experts
- Lepikhin et al. (2020) — GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding