Skip to content
Menu

Notes ·

896 Experts, 16 at a Time

In Kimi K3: Open Frontier Intelligence, Moonshot AI explains how it has pushed mixture-of-experts architecture much further.

“We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts.”

The concept is not new. Mixtral selects 2 of 8 experts, DBRX selects 4 of 16, DeepSeek-V3 selects 8 of 256 routed experts, and Qwen3-Coder selects 8 of 160. K3 scales that to 896 experts while using only 16 for each token.

What may be unique is the extreme sparsity at 2.8 trillion parameters, along with new routing and balancing techniques designed to keep all of those experts useful.

That is more interesting to me than another benchmark ranking. I want to know whether that many experts develop distinct specialties and how reliably the router can choose the right few for the work.

Read the full Kimi K3 announcement.

All notes