AI Dynamics

Global AI News Aggregator

About

Why focus on inference engines: performance gains with vLLM and Sglang

Why do I focus on Inference Engines/Software Stacks for your hardware? – 2x RTX 3090s: ~14.5 tok/s → ~64 tok/s moving to vLLM w/ TP=2 – RTX PRO 6000: ~32 tok/s → ~110 tok/s moving to Sglang So: – CUDA/2+ GPUs: ExLlamaV3/vLLM/Sglang > llama.cpp – Edge: llama.cpp > Ollama

→ View original post on X — @theahmadosman