(1/n) Paper drop: https://
arxiv.org/abs/2405.15743 TLDR: We introduce the sparse maximal update parameterization (SμPar), which ensures optimal HPs remain the same for any width or sparsity level. This dramatically reduces HP tuning costs, allowing SμPar to achieve superior losses.
Sparse Maximal Update Parameterization Reduces Hyperparameter Tuning
By
–
