I'm still a bit surprised that you think the system (DistBelief) that first combined large scale model- and data- parallelism with backprop-based training of very large neural nets is a "dead end", given that nearly every large-scale model that is trained today combines these
DistBelief Legacy: Model and Data Parallelism in Modern Neural Networks
By
–