> E.g. Llama 3 405B used 30.8M GPU-hours, while DeepSeek-V3 looks to be a stronger model at only 2.8M GPU-hours (~11X less compute). Super interesting! And DeepSeek was trained in H800’s which are probably also a tad (or noticeably?) slower than Meta’s H100’s.
DeepSeek-V3 achieves superior performance with 11x less compute than Llama
By
–