AI Dynamics

Global AI News Aggregator

About

DFlash boosts inference 15x on NVIDIA Blackwell with low latency

Increase inference performance by up to 15x without sacrificing responsiveness. DFlash, an open source lightweight block diffusion model designed for speculative decoding, delivers up to 15x higher throughput on NVIDIA Blackwell while maintaining the same user interactivity

→ View original post on X — @nvidiaai