JetMoE
Free while signed in. Answers cite the passages they came from.

MyShell's JetMoE-8B is an open MoE model trained for under $100K that matches or beats LLaMA2-7B, showing that competitive LLM training can be achieved on modest budgets with public data.
MoA + MoE architecture: 24 blocks each combine a Mixture of Attention heads (MoA) and a Mixture of MLP Experts (MoE), with 8 experts per layer and top-2 activation giving 2.2B active parameters out of 8B total.
Cheap training: Trained on 1.25T publicly available tokens using 96 H100s for about two weeks at under $100K of compute, an order of magnitude less than typical 7B training runs.
Benchmark wins: Beats LLaMA2-7B on MMLU (49.2 vs 46.9) and GSM8K (27.8 vs 14.5), and is competitive with far larger open models on standard benchmarks.
Democratization signal: Full weights and training recipe are released, reinforcing that efficient MoE + strong data pipelines let smaller labs ship frontier-quality open models.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack