As @MosaicML showcased in April, on GPT training H100 is ~3x the speed of A100, if you use FP8 training, which is both seriously impressive and grounds you in relative improvement versus the previous generation! https://
mosaicml.com/blog/coreweave
-nvidia-h100-part-1
…
H100 Training Speed Triples A100 with FP8 Optimization
By
–