AI Dynamics

Global AI News Aggregator

About

Scaling AI Training Across Thousands of H100 GPUs

There's three parts. 1. Fitting as large of a network and as large of a batch-size as possible onto the 10k/100k/1m H100s — parallelizing and using memory-saving tricks.
2. Communicating state between these GPUs as quickly as possible
3. Recovering from failures (hardware,

→ View original post on X — @soumithchintala