New from DeepSeek: Mega MoE! Instead of running MoE as a chain of separate steps (dispatch → MLP → combine), Mega MoE fuses everything into a single mega-kernel. Even more importantly, it overlaps NVLink communication with Tensor Core computation, reducing the classic
DeepSeek Mega MoE: Fusing MoE Operations Into Single Kernel
By
–
