AI Dynamics

Global AI News Aggregator

About

W4A8 Inference Production-Ready Integration in vLLM

Excited to share our work on production-ready W4A8 inference, now integrated in vLLM! By combining 4-bit weights (low memory) with 8-bit activations (high compute), we hit the sweet spot for both decoding and prefill — up to 58% faster TTFT and 45% faster TPOT vs W4A16 on Hopper.

→ View original post on X — @cohere