AI Dynamics

Global AI News Aggregator

About

DFlash: Drop-in Speculative Decoding for SGLang, vLLM, TensorRT-LLM

/7 Drop-in for SGLang, vLLM, and TensorRT-LLM. No code refactoring. SGLang:
–speculative-algorithm DFLASH
–speculative-draft-model-path z-lab/Qwen3-8B-DFlash-b16 vLLM: via the Speculators library (
http://
docs.vllm.ai/projects/specu
lators
…, algorithm "dflash") MIT license. ICML 2026 accepted.

→ View original post on X — @alphasignalai