Llama 70B uses multiple WSEs 🙂
@cerebras
-
Cerebras Announces 3x Faster Inference Speed Update
By
–
After this huge speed update we will be focusing on supporting additional customer models, context, and capacity. Stay tuned for more updates!
Chat: http://
Inference.cerebras.ai
API key: http://
cloud.cerebras.ai
Blog: https://
cerebras.ai/blog/cerebras-
inference-3x-faster
… -
Cerebras Optimizes Inference Engine Across Wafer Architecture
By
–
The first release of Cerebras Inference utilized only a fraction of the wafer’s bandwidth, compute, and IO capacity. With this release, we’ve re-written and optimized everything from kernels (matmul, broadcast/reduce) to ML (speculative decoding). Model is still 16-bit and the
-
Cerebras Powers Instant Inference for Major AI Applications
By
–
Numerous companies are using Cerebras to make their inference run at instant speed. These include:
– @GSK for drug discovery
– @Livekit for voice AI
– @tavus for digital twins
– @vellum_ai for testing & iteration -

Cerebras Leads in First Token Latency with Wafer-Scale Integration
By
–
Time to first token is critical for real time applications. Cerebras is among the fastest in first token latency, showing the advantage of wafer scale integration vs. complex networked solutions.
-

Cerebras Wafer Scale Engine runs Llama 70B 184x faster
By
–
Cerebras Inference running Llama 70B is now so fast that it outruns GPU based inference running Llama 3B. The Wafer Scale Engine runs a model 23x larger and 8x faster for a combined 184x performance gain.
-

Cerebras Triples Inference Speed to 2100 Tokens per Second
By
–
We broke all records when we launched Cerebras Inference in August. Today we are tripling our performance from 650 t/s to 2100 t/s.
Cerebras Inference speed is in a league of its own – 16x faster than the fastest GPU solution, 68x faster than hyperscale clouds, and 4-8x faster -
Cerebras Inference 3x Faster: Llama 70B Reaches 2,100 Tokens/Second
By
–
🚨 Cerebras Inference is now 3x faster:
— Cerebras (@cerebras) 24 octobre 2024
Llama3.1-70B just broke 2,100 tokens/s
– 16x faster than the fastest GPU solution
– 8x faster than GPUs running Llama *3B*
– It's like the perf of a new hardware generation in a single software release
Available now at… pic.twitter.com/9VgGWGO6qYCerebras Inference is now 3x faster: Llama3.1-70B just broke 2,100 tokens/s
– 16x faster than the fastest GPU solution
– 8x faster than GPUs running Llama *3B*
– It's like the perf of a new hardware generation in a single software release
Available now at -

Cerebras CrewAI NYC Hack Night Fast Inference Agents
By
–
Come hack with us! Join Cerebras and @crewAIInc for an NYC hack night next Wed, Oct 30. Explore what you can build with fast inference and AI agents RSVP here: https://
lu.ma/92uz16l4 -
Creating Training Data for Llama: Scale and Meta Partnership
By
–
Creating the training data for Llama
— Cerebras (@cerebras) 4 octobre 2024
In this fireside interview, Sophia Luo, a current Greylock Partner and previous Scale Senior PM shares on the long-going Scale and Meta partnerships. She discusses the challenges around creating data to train LLMs like llama, and how this was… pic.twitter.com/adTiAnkQL1Creating the training data for Llama In this fireside interview, Sophia Luo, a current Greylock Partner and previous Scale Senior PM shares on the long-going Scale and Meta partnerships. She discusses the challenges around creating data to train LLMs like llama, and how this was