So GPT-5 is running on GB200, to take full advantage of it, it could be in FP4. So I wonder if this change is hiding the size of the model somewhat. If it was FP8 and on H100, it would be now running at something like 10 tokens per second rather than current 50 tps.
GPT-5 GB200 FP4 Precision and Token Speed Analysis
By
–