Thank you @theinformation
! We're focused on one thing: building the fastest, most efficient infrastructure for AI inference. As agentic AI scales, we believe premium inference will define the next era of computing.
@sambanovaai
-
SambaNova thanks The Information, focuses on fastest AI inference
By
–
-
Heterogeneous disaggregated inference is the future of AI
By
–
Training builds models. Inference builds businesses. At @deeptechweek SF, our Chief Product & Strategy Officer Abhi Ingle shared why heterogeneous, disaggregated inference is the future of AI, and why "more intelligence per joule" is the metric that matters. @LipBuTan1
-

SambaNova’s fastest inference cloud and Ricoh’s custom Japanese AI models
By
–

Customers Scaling AI Faster General Compute @fastinference launched the world's fastest inference cloud for AI agents, powered by SambaNova. Meanwhile, @ricoh is using SambaCloud for custom Japanese AI models and agentic business workflows, moving from tens of tokens per
-
Build Faster Coding Agents with SambaNova Responses API and MiniMax M2.7
By
–
Build Faster Coding Agents SambaNova now supports the Responses API, giving AI engineers a cleaner way to connect modern coding agents to fast, production-ready models. Pair the Responses API with @MiniMax_AI M2.7 for repo edits, tests, patches, and high-volume coding
-
SambaNova demonstrates first disaggregated inference cloud for AI agents
By
–
Disaggregated Inference Is Live! At COMPUTEX, we demonstrated the world's first disaggregated inference cloud for AI agents.
GPUs for prefill. RDUs for decode. CPUs for orchestration. The result: faster agent workflows and better economics than homogeneous infrastructure. -
SambaNova CEO on Vista Equity, Cambium Networks cloud initiative and inference battleground
By
–
Our CEO @RodrigoLiang & Monti Saroya joined @theinformation TITV today to discuss @Vista_Equity and @CambiumNetworks
's cloud initiative & why inference is becoming the next battleground in AI. Disaggregated inference is how we get the speed, efficiency, & economics needed for -
AI Inference: Combining GPU, RDU and CPU for Each Task
By
–
One chip to rule them all? No thanks.
— SambaNova (@SambaNovaAI) 18 juin 2026
AI inference isn't one workload, it's a series of different jobs. That's why you pair GPUs, RDUs & CPUs together, letting each do what it does best.
Better speed, better performance, better economics. 🦾
Chat with us at @RaiseSummit:… pic.twitter.com/Kq5dK9fKtZ—
One chip to rule them all? No thanks. AI inference is not a single workload, it’s a series of different tasks. That’s why we combine GPU, RDU, and CPU together, letting each do what it does best. Better speed, better
— -
SambaNova ready for RaiseSummit 2026: inference 2.0 with GPU and RDU
By
–
We're ready for @RaiseSummit 2026 🙌
— SambaNova (@SambaNovaAI) 16 juin 2026
Training built the models. Inference put them to work. Inference 2.0 is about disaggregating the workload.
GPUs for prefill. RDUs for decode. The right chip for the right job.
Join us at RAISE Summit in Paris to see what's next:… pic.twitter.com/FNThE5nXi8We are ready for @RaiseSummit 2026 Training built the models. Inference put them to work. Inference 2.0 is about disaggregating the workload. GPU for prefilling. RDU for decoding. The right chip for the right task.
-

SambaNova’s disaggregated inference cuts agent latency by 48%
By
–
Same model. Same workload. Different architecture. According to @ArtificialAnlys
, disaggregating prefill and decode reduced agent trajectory latency from 310 seconds to 162 seconds. The right chip for the right workload changes everything. Read more: https://
sambanova.ai/blog/first-dis
aggregated-inference-demo-for-ai-agents-live?utm_source=x&utm_medium=organic
… -
Disaggregated Inference: GPU for Prefill, RDU for Decode
By
–
AI agents spend their time doing two very different jobs: understanding context and generating responses.
— SambaNova (@SambaNovaAI) 15 juin 2026
Disaggregated inference sends each stage to the hardware best built for it.
GPUs for prefill. RDUs for decode. Better performance from both. ⚡ pic.twitter.com/KFXRdUBLWxAI agents spend their time performing two very different tasks: understanding context and generating responses. Disaggregated inference sends each step to the hardware best suited for that task. GPU for prefill. RDU for decode. Better.