Cerebras is proud to be recognized on the 2023 #XB100 list of the top 100 private deep tech companies in the world! Thank you for the recognition @BessemerVP @XPRIZE You can check out the full list here: https://
hubs.li/Q01Tn3J50
@cerebras
-

Cerebras Recognized Top 100 Private Deep Tech Company
By
–
-

SlimPajama Reduces AI Training Compute Costs by 50 Percent
By
–
Why we built SlimPajama – it's all about training efficiency. Without de-duplication, a model would have to go through 1.2T tokens before seeing ~600B unique tokens. SlimPajama sees 600B tokens in half the time. That saves you 50% on compute costs! https://
cerebras.net/blog/slimpajam
a-a-627b-token-cleaned-and-deduplicated-version-of-redpajama
… -
Cerebras Andromeda Supercomputer AI Cluster Infrastructure
By
–
There are a lot of galaxies in the night sky. Did you really have to pick the only one already used for an AI cluster https://
venturebeat.com/ai/cerebrass-a
ndromeda-supercomputer-has-13-5m-cores-that-can-do-an-exaflop-in-ai-computing/
… -
SlimPajama Dataset: Preprocessing Library for LLM Training
By
–
SlimPajama dataset – https://
lnkd.in/gCchZ-xz
Preprocessing library: https://
lnkd.in/gV7r3YNC
Read our blog: -

SlimPajama: Open-Source Cleaned RedPajama Dataset Released
By
–
We recently announced the availablity of SlimPajama – an open-source, cleaned, and deduplicated version of RedPajama-1T. It is half the size and trains twice as fast and when upsampled, performs equal or better than RedPajama. See below for the dataset and preprocessing library
-
SlimPajama Tools Released for Building AI Models
By
–
Here are the tools we created to build SlimPajama. Have fun! x.com/dmsobol/status…
-
SlimPajama Tools Released for Open Source AI
By
–
Here are the tools we created to build SlimPajama. Have fun!
-
RedPajama Project Launches with OpenTensor and Together Support
By
–
We’d like to thank our partner @opentensor for supporting this project. And credit goes to @togethercompute and the entire team that created the RedPajama dataset! We can’t wait to see what you’ll build. Join our Discord and let us know your feedback: https://
discord.com/channels/10859
60591052644463/1085960592050896937
… -
Custom Parallel Data Pipeline for Trillion Token Deduplication
By
–
It was no mean feat to deduplicate data on this scale – existing tools does not scale to a trillion tokens. We built a custom parallel data pre-processing pipeline and are sharing the code open source with the community.
-

SlimPajama: 50% Smaller, Twice Faster LLM Training Dataset
By
–
SlimPajama cleans and deduplicates RedPajama-1T, reducing the total token count and file size by 50%. It's half the size and trains twice as fast! It’s the highest quality dataset when training to 600B tokens and when upsampled performs equal or better than RedPajama.