(7/n) We test Sparse Iso-Parameter Scaling, where width and sparsity are increased, holding n_params constant. Using the standard practice (SP + dense optimal HPs), one would conclude only up to 50% sparsity help. Instead using SμPar allows an 87.5% sparse model to match dense
@cerebras
-

SμPar Extends μP for Sparse Neural Network Training
By
–
(4/n) We introduce the sparse maximal update parameterization (SμPar), which extends μP to also fix the vanishing activation problem with increasing sparsity.
-

SμPar Maintains Stable Hyperparameters Across Sparsity Width
By
–
(5/n) SμPar enables optimum hyperparameters to remain stable for any combination of sparsity and width, unlike SP and μP.
-

Preventing Activation Vanishing in Sparse Neural Network Training
By
–
(3/n) This motivated us to develop a parameterization that prevents activation vanishing due to sparsity. We need to ensure all 3 operations in a training step (forward, backward, and weight update) are controlled with respect to sparsity.
-

Sparse Maximal Update Parameterization Reduces Hyperparameter Tuning
By
–
(1/n) Paper drop: https://
arxiv.org/abs/2405.15743 TLDR: We introduce the sparse maximal update parameterization (SμPar), which ensures optimal HPs remain the same for any width or sparsity level. This dramatically reduces HP tuning costs, allowing SμPar to achieve superior losses. -

Weight Sparsity Impact on Activation Scales Training
By
–
(2/n) Increasing weight sparsity causes vanishing activation scales with both SP and μP, leading to poor training dynamics.
-

Company Named to TIME 100 Most Influential Companies List
By
–
We are proud and honored today to have been named to the TIME 100 Most Influential Companies list! As we reflect on this incredible moment, we are grateful for the fundamentals that got us here: Winning team – We have the very best engineers and talent worldwide,
-

AI Reshaping Healthcare: Cerebras at GenAI Summit SF 2024
By
–
Learn about AI in Healthcare! VP and Field CTO Natalia Vassilieva will be speaking at the #GenAISummitSF2024 on "How AI is Reshaping the Service in Medical and Healthcare" Register here: http://
genaisummit.ai Learn how Cerebras is reshaping Healthcare: -

Cerebras CEO Breaks GPU Barriers for AI Model Training
By
–
Check out our CEO Andrew Feldman’s keynote on how Cerebras has broken through GPU barriers and makes advanced AI model training dead simple at the Mint Digital Innovation Summit 2024. Watch Andrew’s talk: https://
livemint.com/industry/ai-is
-easy-to-describe-challenging-to-do-cerebras-systems-ceo-andrew-feldman-11716561044183.html
… Contact Cerebras to accelerate your AI -

MediSwift Biomedical Language Models Accepted to ACL 2024
By
–
MediSwift has been accepted into ACL 2024! MediSwift is the first suite of biomedical language models that employ sparse pre-training techniques to significantly reduce computational costs while outperforming existing models up to 7B parameters on benchmark tasks such as
