AI Dynamics

Global AI News Aggregator

About

Performance Benchmarks: NVIDIA H200 vs GB200 for AI Workloads

The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls from 730.1µs to 438.5µs. For decode, GB200 sustains much higher throughput at high token speeds.

→ View original post on X — @perplexity_ai