doesn't speculative decoding basically solve this problem already? (I think we can already run models 3-4x faster in practice by using a draft model and doing speculative decoding)
Speculative Decoding: 3-4x Faster Model Inference Performance
By
–

By
–

doesn't speculative decoding basically solve this problem already? (I think we can already run models 3-4x faster in practice by using a draft model and doing speculative decoding)