AI Dynamics

Global AI News Aggregator

About

Optimizing Throughput for Large MoE Models on GB 200 Hardware

GB 200s change how one does the prefill and decode disaggregation when serving large MoEs like Qwen. We’ve published details of our stack quantifying the throughput benefits compared to serving on Hoppers.

→ View original post on X — @aravsrinivas