I feel the need, the need for speed! Llama 4 Maverick from @AIatMeta is now available on SambaNova Cloud & it's the fastest inference verified by @ArtificialAnlys at 655 t/s. $0.50 / million input tokens & $2.00 per million output tokens. On SambaNova Cloud
@sambanovaai
-

Llama4 Fastest Independent Analysis Output Speed Results
By
–
The GOAT of independent analysis has spoken @ArtificialAnlys Co-Founder & CEO, @_micah_h shares their findings on our output speed. You heard it straight from the source: we're the FASTEST! #Llama4
-

SambaNova Launches Llama 4 Scout at Record Inference Speed
By
–
Herding in a new era of the FASTEST AI inference As an official @AIatMeta launch partner, we're bringing #Llama4 models to life on SambaNova Cloud, powered by RDUs! Llama 4 Scout is available now, running at 697 t/s Llama 4 Maverick coming soon Read more
-

Llama 4 Scout Available on SambaNova Cloud at 697 Tokens/Second
By
–
Llama 4 Scout from @AIatMeta is now available on SambaNova Cloud! Fastest Inference clocked at 697+t/s. Llama 4 Maverick will be out next week, followed by higher context lengths up to 128K! Try it now on SambaNova Cloud
-
Llama 4 Scout Achieves Record 697 Tokens Per Second Speed
By
–
Don’t blink! You might miss just how fast we’re going⚡️
— SambaNova (@SambaNovaAI) 7 avril 2025
🚀 697 t/s on @AIatMeta's #Llama4, independently verified by @ArtificialAnlys
"…the fastest output speed we have measured yet for Llama 4 Scout.” — @_micah_h
Try it now 👇Don’t blink! You might miss just how fast we’re going 697 t/s on @AIatMeta
's #Llama4, independently verified by @ArtificialAnlys "…the fastest output speed we have measured yet for Llama 4 Scout.” — @_micah_h Try it now -

Claude and Llama Struggle with Japanese Document Understanding
By
–
Claude 3.5/Llama 3.2 ace English DocVQA, but how do they fare in Japanese? New benchmark alert: JDocQA (curated by @NAIST_MAIN
, @RIKEN_RCCS
, ATR) exposes multilingual gaps in top VLMs. Read more -
Meta and SambaNova Partner to Launch Llama 4 Models
By
–
We've teamed up with @AIatMeta to unleash the power of both Llama 4 models! 🦙⚡️
— SambaNova (@SambaNovaAI) 5 avril 2025
Get ready for lightning-fast AI magic—dropping soon on SambaNova Cloud. Stay tuned, the future’s speeding up! 🚀 pic.twitter.com/44fnqCmbFOWe've teamed up with @AIatMeta to unleash the power of both Llama 4 models! Get ready for lightning-fast AI magic—dropping soon on SambaNova Cloud. Stay tuned, the future’s speeding up!
-
Meta Launches Llama 4 on SambaNova Cloud Platform
By
–
We are excited to be partnering with @AIatMeta to launch Llama 4. Available soon on SambaNova Cloud with blazing fast inference!
-
DeepSeek Model 4x Faster: One Rack Versus Forty
By
–
🔥 “We’re almost 4x faster than anyone else on the @deepseek_ai model, on one rack vs 40 racks.” @RodrigoLiang speaking to @Bloomberg's @mattmiller1973 at yesterday's event in NYC.
— SambaNova (@SambaNovaAI) 3 avril 2025
🦾 No fluff, just compute power that actually keeps up with Gen AI’s demands. #DeepSeek pic.twitter.com/jR15TNB8US“We’re almost 4x faster than anyone else on the @deepseek_ai model, on one rack vs 40 racks.” @RodrigoLiang speaking to @Bloomberg
's @mattmiller1973 at yesterday's event in NYC. No fluff, just compute power that actually keeps up with Gen AI’s demands. #DeepSeek -

SambaNova Cloud Achieves Record 250 Tokens per Second Inference
By
–
Breaking speed limits left & right with #AI inference! Our high-speed support for @deepseek_ai R1 671B delivers 250 t/s per user, leaving other GPU-powered solutions in the dust Devs — get ready to turbocharge your AI with SambaNova Cloud More on our blog
