ICYMI: We've partnered with @HuggingFace to bring 10x faster #AI inference speeds to devs! @AI at Meta's Llama 3 & @Alibaba_Qwen models SambaNova Cloud's Llama Guard & Qwen QwQ Easy integration with minimal code changes @deepseek_ai coming soon! Try it now
LLMS
-
DeepSeek Shifts: China Advances, Open Models Commoditize AI
By
–
The buzz over DeepSeek this week crystallized, for many people, a few important trends that have been happening in plain sight: (i) China is catching up to the U.S. in generative AI, with implications for the AI supply chain. (ii) Open weight models are commoditizing the
-
Bloomberg segment on AI and DeepSeek developments
By
–
I got on Bloomberg News yesterday to talk about AI and the DeepSeek hullabaloo. You can find my segment starting on minute 58.
-
Mistral Small 3 Available on Partner Platforms
By
–
In addition to being available on la Plateforme (
http://
console.mistral.ai), Mistral Small 3 is available on our partner platforms -

Teaching LLMs with Textbook Information Structure
By
–
We have to take the LLMs to school. When you open any textbook, you'll see three major types of information: 1. Background information / exposition. The meat of the textbook that explains concepts. As you attend over it, your brain is training on that data. This is equivalent
-
Cerebras launches world’s fastest DeepSeek R1 Llama-70B inference
By
–
Cerebras DeepSeek R1 Llama-70B is available now on http://
cerebras.ai, and we're offering API preview access to select customers. Read our blog to learn more. https://
cerebras.ai/blog/cerebras-
launches-worlds-fastest-deepseek-r1-llama-70b-inference
… -

Cerebras Achieves 1,500 Tokens/Sec with R1 70B Model
By
–
Cerebras is the fastest reasoning platform in the world. We run R1 70B at over 1,500 tokens/s – 57x faster than GPU solutions.
-
Cerebras Inference Makes Reasoning Models Instant
By
–
Reasoning models are powerful but can take minutes to generate the final answer. Cerebras Inference makes AI instant again. In the coding example below, Cerebras Inference R1 70B returns the answer in 1.5 seconds vs. 22 seconds using o1 mini. pic.twitter.com/43TQ9YVl8l
— Cerebras (@cerebras) 30 janvier 2025Reasoning models are powerful but can take minutes to generate the final answer. Cerebras Inference makes AI instant again. In the coding example below, Cerebras Inference R1 70B returns the answer in 1.5 seconds vs. 22 seconds using o1 mini.
-

DeepSeek R1 70B Outperforms GPT-4o and o1-mini
By
–
DeepSeek’s R1 70B combines the powerful reasoning ability of the full R1 model with the size and speed of Llama 70B. R1 70B outperforms GPT-4o and o1-mini across a range of general and reasoning benchmarks, making it the most capable Llama 70B variant by far.
-

DeepSeek R1 70B Now Available on Cerebras Infrastructure
By
–
DeepSeek R1 70B is now on Cerebras!
– Instant reasoning at 1,500 tokens/s – 57x faster than GPUs
– Higher model accuracy than GPT-4o and o1-mini
– Runs 100% on Cerebras US data centers https://
inference.cerebras.ai
