Algorithm 1) All-reduce
— Akshay 🚀 (@akshay_pachaar) 17 août 2025
An obvious way is to send the gradients from one device to all other devices to synchronize them.
But this utilizes high bandwidth.
If every GPU has “N” elements and there are “G” GPUs, it results in a total transfer of G*(G-1)*N elements 👇 pic.twitter.com/cxqdc3PszF
Algorithm 1) All-reduce An obvious way is to send the gradients from one device to all other devices to synchronize them. But this utilizes high bandwidth. If every GPU has “N” elements and there are “G” GPUs, it results in a total transfer of G*(G-1)*N elements
