AI Dynamics

Global AI News Aggregator

About

Speculative Decoding: 3-4x Faster Model Inference Performance

doesn't speculative decoding basically solve this problem already? (I think we can already run models 3-4x faster in practice by using a draft model and doing speculative decoding)

→ View original post on X — @jxmnop