AI agents spend their time doing two very different jobs: understanding context and generating responses.
— SambaNova (@SambaNovaAI) 15 juin 2026
Disaggregated inference sends each stage to the hardware best built for it.
GPUs for prefill. RDUs for decode. Better performance from both. ⚡ pic.twitter.com/KFXRdUBLWx
AI agents spend their time performing two very different tasks: understanding context and generating responses. Disaggregated inference sends each step to the hardware best suited for that task. GPU for prefill. RDU for decode. Better.