/5 The insight is from Samragh et al. (2025): large autoregressive LLMs already encode information about multiple future tokens in their hidden states. The target model is doing work that the drafter never gets to see. DFlash taps that. It extracts hidden features from uniformly
DFlash extracts hidden features from future tokens in LLMs
By
–