Same prompt. Same model. Two stacks.
— SambaNova (@SambaNovaAI) 9 juin 2026
At #Computex, we demonstrated disaggregated inference live: GPUs handling prefill, SambaNova RDUs handling decode, and CPUs orchestrating agent execution.
The result? Up to 2X the speed of B200-only configurations 🦾 pic.twitter.com/YYP8o6WYrK
Same prompt. Same model. Two stacks. At #Computex, we demonstrated disaggregated inference live: GPUs handling prefill, SambaNova RDUs handling decode, and CPUs orchestrating agent execution. The result? Up to 2X the speed of B200-only configurations








