(4/n) We introduce the sparse maximal update parameterization (SμPar), which extends μP to also fix the vanishing activation problem with increasing sparsity.
SμPar Extends μP for Sparse Neural Network Training
By
–

By
–

(4/n) We introduce the sparse maximal update parameterization (SμPar), which extends μP to also fix the vanishing activation problem with increasing sparsity.