7). MagicDec – shows how speculative decoding can enhance throughput, reduce latency, and maintain accuracy in long context generation scenarios; it finds that as sequence length and batch size increase, bottlenecks shift from compute-bound to memory-bound…
MagicDec: Speculative Decoding Enhances LLM Throughput and Latency
By
–