Distbelief was used internally for training 10,000+ distinct models of all kinds of architectures (convnets, RNNs, LSTMs, feed forward networks, sparse MoEs, etc), with all kinds of different training objectives (supervised, unsupervised, RL, etc).
DistBelief trained 10000+ models across architectures
By
–