MoE Training, Part 1 — in one tweet:
— Satya Mallick (@LearnOpenCV) 21 mai 2026
You do NOT assign "this expert handles medicine, this one handles law."
You start with 9 random experts + a router. The router learns to pick 2–3 per question. Specialization emerges from data, not design.
That's how Mixtral and DeepSeek… pic.twitter.com/YGSc6ak94i
MoE Training, Part 1 — in one tweet:
You do NOT assign "this expert handles medicine, this one handles law."
You start with 9 random experts + a router. The router learns to pick 2–3 per question. Specialization emerges from data, not design.
That's how Mixtral and DeepSeek