AI Dynamics

Global AI News Aggregator

About

SμPar Extends μP for Sparse Neural Network Training

(4/n) We introduce the sparse maximal update parameterization (SμPar), which extends μP to also fix the vanishing activation problem with increasing sparsity.

→ View original post on X — @cerebras