It turns out you can drop 80% lowest-entropy tokens and still get good performance. This paper shows that only ~20% of tokens—those with the highest entropy—actually drive the reasoning process in LLMs trained with Reinforcement Learning with Verifiable Rewards (RLVR). The rest?
High-Entropy Tokens Drive LLM Reasoning in RLVR
By
–
