and confusingly, we can predict the next-next token from hidden states at a given timestep. it just turns out that this is seems to be mostly a property of language — for example if you say "New" then it's somewhat likely that "City" is two tokens away
Language Models Predict Next-Next Tokens From Hidden States
By
–