We will update more Gemma 4 31B information as we go! https://
inference-docs.cerebras.ai/models/gemma-4
-31b
…
@cerebras
-
Cerebras to share more Gemma 4 31B information updates
By
–
-
Homomorphic encryption for the privacy of LLM queries
By
–
Right now, when you send a query to an LLM, it gets decrypted on the server. The LLM sees your data in plain text.
— Cerebras (@cerebras) 19 juin 2026
Prof. Ajay Joshi (BU, @CipherSonicAI ) on fully homomorphic encryption, which may be key for the future of AI privacy: how we can compute on data without ever… pic.twitter.com/vCfcUGRVLjCurrently, when you send a query to an LLM, it is decrypted on the server. The LLM sees your data in plain text. Prof. Ajay Joshi (BU, @CipherSonicAI) on fully homomorphic encryption, which could be the key to the future of privacy of
-

Early Access to Gemma 4 on Cerebras and 24h Hackathon with $5000
By
–
Early access to the world's fastest multimodal model is available. Get your hands on the Gemma 4 model on Cerebras. 24-hour hackathon with a prize of $5,000 and the flagship project presented by Google DeepMind and Cerebras. RSVP link in comments
-

Fast inference: more context, tools, validation, and controls
By
–
Faster inference offers cybersecurity products much more than just faster responses. It allows them to integrate more context, tool calls, validation, and policy controls within their latency budget. More details in our blog.
-
Andrew Feldman compares AI selection to shopping at Costco
By
–
Andrew Feldman on AI model selection:
— Cerebras (@cerebras) 15 juin 2026
Shop like you're at Costco.
You don't need to buy everything.
You don't need the biggest thing on the shelf.
And somehow you definitely don't need the 400 oz tub of mayo.
Same with AI.
Not every problem needs the biggest, most expensive… pic.twitter.com/WxC3bwXn8TAndrew Feldman on AI model selection: Do your shopping as if you were at Costco. You don't need to buy everything. You don't need the biggest thing on the shelf. And somehow, you certainly don't need the pot.
-
NVIDIA and AWS bet on disaggregated inference for AI infrastructure
By
–
NVIDIA paid $20B for Groq.
— Cerebras (@cerebras) 11 juin 2026
AWS partnered with Cerebras for the same purpose.
A quick breakdown of why disaggregated inference is the next thing in AI infrastructure. pic.twitter.com/kP8mDpPl9cNVIDIA paid $20 billion for Groq.
AWS partnered with Cerebras for the same purpose.
A brief overview of why disaggregated inference is the next step in AI infrastructure. -

Speed comparison between Gemini 3.5 Flash and Kimi K2.6 on Cerebras
By
–
Google just released its fastest model: Gemini 3.5 Flash. We directly pitted it against Kimi K2.6 on Cerebras. Both are equal in intelligence, but what about speed? Full benchmark results: https://cerebras.ai/blog/which-is-faster-gemini-3-5-flash-or-kimi-k2-6-on-cerebras
… -
We are in the recursive era of AI according to Cerebras
By
–
We are in the recursive era of AI. Cerebras wrote an article about this from a hardware perspective in March: https://cerebras.ai/blog/why-the-ai-race-shifted-to-speed …
-
The Recursive Era of AI: Fast Inference and Accelerated Development per Cerebras
By
–
We are in the recursive era of AI.
Faster inference => faster AI development.
Our take on this from a hardware perspective:
https://cerebras.ai/blog/why-the-ai-race-shifted-to-speed
… -
Colocating memory avoids costly transfers during inference
By
–
"If you can co-locate your memory, you're getting a lot more bang for your buck because you're avoiding this costly memory transfer."@sarahookr (author of The Hardware Lottery, founder of @adaptionlabs ) on why inference is forcing a new chip paradigm – one that wafer-scale was… pic.twitter.com/tTfu19zUWU
— Cerebras (@cerebras) 4 juin 2026“If you can colocate your memory, you get much more value for your money because you avoid that costly memory transfer.” @sarahookr (author of The Hardware Lottery, founder of @adaptionlabs) on why inference forces a new paradigm