In a recent interview with India's @EconomicTimes
, Cerebras CEO Andrew Feldman said he categorically considers India as a priority for Cerebras, citing the country's tremendous engineering talent, top universities, and a growing AI ecosystem. Read here:
@cerebras
-
Cerebras Prioritizes India for AI Hardware Development
By
–
-
Cerebras Nominated for HPC Industry Awards and New Products
By
–
(2/2) Please vote for Cerebras and our partners in the following categories: -Best Use of HPC in Industry
-Best HPC Collaboration
-Top Supercomputing Achievement
-Top 5 New Products or Technologies to Watch
-Top 5 Vendors to Watch -
Cerebras Nominated for HPCwire Readers’ Choice Awards
By
–
(1/2) ICYMI, Cerebras and our partners have been nominated for the @HPCwire Readers’ Choice Awards in various categories! Please vote for Cerebras and our partners in the following categories by October 1 by clicking here: https://
hpcwire.com/2023-hpcwire-r
eaders-choice-awards-voting-is-open/
… -

BTLM-3B Outperforms Larger 7B Models in Long Context
By
–
The paper explores long context performance in detail. We found BTLM-3B-8K outperforms other 7B-8K models despite being trained with less than a fifth the pretraining compute and less than half the size.
-

ALiBi Enables Context Extrapolation Beyond Training Length
By
–
ALiBi allows our model trained on 8192 context length to extrapolate well up to 9216 tokens out of the box. We also show that ALiBi models only achieve near-perfect extrapolation when they are severely undertrained (<1 tokens/parameter).
-

BTLM: World’s Most Accurate 3B Parameter Open Source Model
By
–
BTLM is the world’s most accurate 3B parameter model, with performance that rivals many 7B models. It’s fully open source and has been downloaded >1M times.
-

LLM Training Techniques: 2.86x Compute Efficiency Gains
By
–
There are so many LLM training tricks but which ones should you use? We quantify the effects of techniques like SwiGLU, ALiBi, etc and show how using them in conjunction can achieve the same loss with 2.86x less pretraining compute or 1.74x less parameters.
-

BTLM-3B-8K: Distilling SOTA LLM Training Recipe
By
–
We just dropped the BTLM-3B-8K paper on arXiv! It distills our recipe for training SOTA LLMs:
– Extensively deduplicated dataset (SlimPajama)
– Hyperparameter search using muP
– Variable sequence length training + ALiBi
– Aggressive LR decay https://
arxiv.org/abs/2309.11568 -

AI Applications in Geothermal Exploration and Carbon Capture
By
–
The Next Platform also mentions, "They also note in the full paper this is also a potential target in areas like geothermal exploration, carbon capture, and storage projects, with greater efficiency and lesser environmental impact." Read the paper here: https://
repository.kaust.edu.sa/handle/10754/6
94388
… -

Cerebras Wafer-Scale Solves Seismic Processing Memory Wall Challenge
By
–
"Cerebras and the KAUST research team are touting this as a solution to the 'memory wall' problem, at least for this domain." @TheNextPlatform wrote about our seismic processing work, which is a finalist for the Gordon Bell Prize Read their story here: https://
nextplatform.com/2023/09/20/sei
smic-data-processing-on-waferscale-has-gordon-bell-prize-potential/
…