"today, Groq is preparing to produce its second-gen chip, which it says will offer a two to three times jump in efficiency across speed, cost, and energy consumption. Ross describes it as 'like skipping from fifth grade all the way to your Ph.D. program.'"
HARDWARE
-
Body Battery Technology Becoming Common Vernacular in SF
By
–
overheard in SF: Is your body battery okay?
-

SambaNova SN40L Chip Breaks World Records for AI Inference
By
–
SambaNova Cloud has broken world records! What's behind our success? Our powerful #AI chip: the SN40L. Read more about why the SN40L is the best #inference solution https://
sambanova.ai/blog/sn40l-chi
p-best-inference-solution
… -
Santosh Raghavan Speaks on AI Chips at Extreme Tech Challenge
By
–
Don't miss Santosh Raghavan, Power Architecture & Silicon Technology Engineer at Groq, speaking on a panel about AI Chips and Datacenter Infrastructure Innovations at Extreme Tech Challenge (XTC) next week in San Francisco.
-

Meta AI Models Deployed on Mobile CPUs with Partner Collaboration
By
–
Thanks to close work with @arm
, @mediatek and @qualcomm
, these new models are ready to deploy on even more mobile CPUs.
We are also currently collaborating with partners to utilize NPUs for these quantized models for even greater performance. -
Cerebras Optimizes Inference Engine Across Wafer Architecture
By
–
The first release of Cerebras Inference utilized only a fraction of the wafer’s bandwidth, compute, and IO capacity. With this release, we’ve re-written and optimized everything from kernels (matmul, broadcast/reduce) to ML (speculative decoding). Model is still 16-bit and the
-
Cerebras Powers Instant Inference for Major AI Applications
By
–
Numerous companies are using Cerebras to make their inference run at instant speed. These include:
– @GSK for drug discovery
– @Livekit for voice AI
– @tavus for digital twins
– @vellum_ai for testing & iteration -
Cerebras Announces 3x Faster Inference Speed Update
By
–
After this huge speed update we will be focusing on supporting additional customer models, context, and capacity. Stay tuned for more updates!
Chat: http://
Inference.cerebras.ai
API key: http://
cloud.cerebras.ai
Blog: https://
cerebras.ai/blog/cerebras-
inference-3x-faster
… -

Cerebras Leads in First Token Latency with Wafer-Scale Integration
By
–
Time to first token is critical for real time applications. Cerebras is among the fastest in first token latency, showing the advantage of wafer scale integration vs. complex networked solutions.
-

Cerebras Wafer Scale Engine runs Llama 70B 184x faster
By
–
Cerebras Inference running Llama 70B is now so fast that it outruns GPU based inference running Llama 3B. The Wafer Scale Engine runs a model 23x larger and 8x faster for a combined 184x performance gain.