It's worth pointing out that we have been pushing on large-scale training and asynchronous techniques for the last ~14 years. Here's our NeurIPS 2012 paper where we demonstrated that this approach could be used to train very large neural networks (for the time: 30X larger than
Large-Scale Asynchronous Training Techniques for Neural Networks
By
–