🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture

OLMoE

First page
OLMoE
Paper summary

introduces a fully-open LLM that leverages sparse Mixture-of-Experts. OLMoE is a 7B parameter model and uses 1B active parameters per input token; there is also an instruction-tuned version that claims to outperform Llama-2-13B-Chat and DeepSeekMoE 16B.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack