Premium inference is powering the next generation of AI agents. First live disaggregated inference demo for AI agents New SN50 RDU purpose-built for agentic inference Faster, more efficient AI with industry-leading throughput See what's next for AI inference:
@sambanovaai
-

World’s first heterogeneous disaggregated inference cloud shown live at ComputeX
By
–
The world's first heterogenous disaggregated inference cloud was just shown running live at ComputeX. VC2 — backed by a $3.5B compute commitment to SambaNova from @Vista_Equity & @cambiumcapital — brings three chips together in production for the first time:
– NVIDIA B200 GPUs -
SambaNova post about Intel’s rack-scale agentic AI design
By
–
Agentic AI needs CPUs, GPUs, and AI accelerators working together. @TheRegister highlights @intel
's new rack-scale agentic AI designs and the first customer deployment of the Intel + SambaNova disaggregated inference blueprint through VC2. The result: GPUs handle prefill, -

SambaNova unveils disaggregated inference demo with 2x speedup
By
–
The first disaggregated inference demo for AI agents is now live. At #COMPUTEX2026, SambaNova demonstrated premium inference running in production at VC2 — using NVIDIA B200 GPUs for prefill and SambaNova RDUs for decode. The result: 2x faster inference than B200-only
-

SambaNova unveils disaggregated inference cloud with Intel Xeon at COMPUTEX2026
By
–
The future of inference is disaggregated. At #COMPUTEX2026, Vector Core Compute unveiled the world's first fully disaggregated inference cloud powered by Intel Xeon and SambaNova RDUs—designed for the scale, speed, and economics modern AI demands. Excited to help bring this
-

First disaggregated inference cloud VectorCore Compute launched at COMPUTEX2026
By
–
At @LipBuTan1
's #COMPUTEX2026 keynote today, @RodrigoLiang stepped onstage with @RFS_Vista to power up the world's first disaggregated inference cloud, VectorCore Compute (VC2), launched by @Vista_Equity and Cambium Capital. Three chips ran disaggregated inference, live from the -
AI agents rely on infrastructure; Ricoh runs Japanese models at 700+ tokens/sec
By
–
AI agents are only as useful as the infrastructure behind them. @ricoh is running custom Japanese AI models on SambaCloud at 700+ tokens/sec, delivering up to 10× faster performance than previous infrastructure. What once took a minute now finishes in ~10 seconds, making
-
CPU Role in Agentic Systems: Inference Orchestration
By
–
In agentic systems, CPUs do two things: orchestrate inference and run everything around it: LLVM compilation, vector DB queries, tool calls.
— SambaNova (@SambaNovaAI) 29 mai 2026
Faster execution at each step = shorter agent loop. That's why Xeon 6 + RDU is the full stack, not just the accelerator. pic.twitter.com/rwYYigDC2EIn agentic systems, CPUs do two things: orchestrate inference and run everything around it: LLVM compilation, vector DB queries, tool calls. Faster execution at each step = shorter agent loop. That's why Xeon 6 + RDU is the full stack, not just the accelerator.
-

Planner-Executor Pattern for Coding Agents
By
–
Not every coding-agent task needs the same model. Use frontier models for planning and reasoning. Use @MiniMax_AI M2.7 on SambaCloud for fast execution, edits, retries, and tool-heavy workflows That’s the planner / executor pattern.
-
General Compute Builds Efficient AI Inference Cloud Infrastructure
By
–
@TimFernholz at @TechCrunch breaks down how General Compute is building its inference cloud with SambaNova, and why faster, more efficient inference infrastructure is becoming critical for the next wave of AI. Featuring insights from General Compute CEO @FPuklowski and CTO
