(3/n) This motivated us to develop a parameterization that prevents activation vanishing due to sparsity. We need to ensure all 3 operations in a training step (forward, backward, and weight update) are controlled with respect to sparsity.
Preventing Activation Vanishing in Sparse Neural Network Training
By
–
