Our VP Product Jessica Liu shared about our new Multimodal model release and the ML expertise that is enabling us to train large models faster than a speeding bullet! Check out our multimodal model checkpoints on Hugging Face! https://
huggingface.co/cerebras
@cerebras
-

Cerebras Releases Multimodal Model Checkpoints on Hugging Face
By
–
-

Cerebras CS-3 Wafer-Scale Architecture Enables Superior Training Performance
By
–
Our CTO and Co-Founder Sean Lie takes us deep into the heart of Cerebras Hardware and our Wafer-Scale Architecture. On CS-3, we are enabling large scale training at orders of magnitude performance advantage compared to GPU. But even our largest clusters natively operate as a
-
Cerebras CEO Discusses AI Hardware Challenges and Large Chip Training
By
–
Our CEO @andrewdfeldman kicked off Cerebras AI Day to a standing room only audience. Andrew’s keynote covered: 1. The AI Capabilities Chasm
2. The GPU Challenge
3. Large Models train best on Large Chips. #AI #AIcompute -

Cerebras AI Day 2026: Hardware Innovation Event Begins
By
–
Cerebras AI Day is today and we can't wait to get this party started! See you in San Jose!
-

MediSwift-XL Sparse Model Outperforms Dense Competitor
By
–
(4/n) At 75% sparsity, MediSwift-XL outperforms the dense MediSwift-Med, despite having the same non-embedding parameters. This highlights the advantages of training larger but sparse models over smaller, densely parameterized models.
-
MediSwift Trained on Cerebras Wafer-Scale Cluster
By
–
(5/n) MediSwift was trained in three sizes (Med, Large, and XL), with dense and sparse variants, on a Cerebras Wafer-Scale Cluster with only a few configuration changes. Cerebras makes it easy to experiment and train production-ready models. Contact us to learn how we can help
-

MediSwift Achieves 75% Sparsity, Reduces Training FLOPs by 2.5x
By
–
(2/n) MediSwift capitalizes on our most recent innovations in sparsity, inducing up to 75% unstructured weight sparsity during in-domain pre-training on biomedical texts. This results in a 2-2.5x reduction in the required training FLOPs. Blog: https://
cerebras.net/blog/sparsity-
made-easy-introducing-the-cerebras-pytorch-sparsity-library
… -

Dense Fine-Tuning Boosts MediSwift Biomedical Task Performance
By
–
(3/n) Dense fine-tuning and soft prompting enhance the performance of sparsely pre-trained MediSwift models on biomedical tasks (e.g., PubMedQA). This ensures high accuracy, thereby improving the efficiency-accuracy Pareto frontier. Pubmed QA leaderboard: https://
pubmedqa.github.io -

MediSwift: Sparse Biomedical Language Models Reduce Computational Costs
By
–
(1/n) Introducing MediSwift, the first suite of biomedical language models that employ sparse pre-training techniques to significantly reduce computational costs, while outperforming existing models up to 7B parameters on benchmark tasks such as PubMedQA. Paper:
-

Cerebras Named America’s Best AI Hardware Startup Employer
By
–
At Cerebras, our people are at the heart of our innovation. We are honored to share that @Forbes and @StatistaCharts have recognized Cerebras as one of America's Best Startup Employers in 2024. https://
forbes.com/lists/americas
-best-startup-employers/?sh=5e1752bd2ad7
… Interested in working at Cerebras? Check out our career