Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Paper: https://
arxiv.org/pdf/2506.01939
.pdf
…
Project: https://
shenzhi-wang.github.io/high-entropy-m
inority-tokens-rlvr/
…
LLMS
-

High-Entropy Minority Tokens Improve LLM Reasoning via Reinforcement Learning
By
–
-

High-Entropy Tokens Drive LLM Reasoning in RLVR
By
–
It turns out you can drop 80% lowest-entropy tokens and still get good performance. This paper shows that only ~20% of tokens—those with the highest entropy—actually drive the reasoning process in LLMs trained with Reinforcement Learning with Verifiable Rewards (RLVR). The rest?
-
LLMs Influencing Parenting Decisions and Family Planning
By
–
Decisions I make in parenting my children are different than they would be if the Internet didn't exist. *Obviously.* Haven't made major decisions due to LLMs on that front but feel like would be almost negligent not to at some point.
-
LLMs Path to Internet-Scale Economic Impact
By
–
(LLMs are, as of today's level of integration into the economy, very short of the impact of the Internet on society, our daily lives, the economy, etc. But I expect them to get there.)
-
LLMs Importance: Lower Bound Internet, Upper Bound Unclear
By
–
I continue to think we're lower bounded on eventually getting to "LLMs are only as important as the Internet", says the guy who thinks the Internet is the magnum opus of the human race. Upper bound: very unclear.
-
LLMs Fundamentally Transform Software Engineering Craft
By
–
I've mentioned that some of the most talented technologists I know are saying LLMs fundamentally change craft of engineering; here's a recently published example from @tqbf
. -
RLHF 101: Technical Tutorial on Reinforcement Learning from Human Feedback
By
–
RLHF 101 from ML@CMU A Technical Tutorial on Reinforcement Learning from Human Feedback "This blog dives into the full training pipeline of the RLHF framework. We will explore every stage — from data generation and reward model inference, to the final training of an
-

AlphaOne: Reasoning Models Thinking Slow and Fast
By
–
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
Paper: https://
arxiv.org/pdf/2505.24863
.pdf
…
Project: https://
alphaone-project.github.io
Code: https://
github.com/ASTRAL-Group/A
lphaOne
… (coming soon) -

AlphaOne: Modulated Reasoning Optimizes LRM Performance
By
–
Large Reasoning Models (LRMs) often struggle to balance deep thinking with efficient generation. Prior methods use monotonic scaling (e.g. "more steps = better results") but lack flexibility. Enter AlphaOne: Modulated Reasoning at Test Time AlphaOne introduces a universal
-
LLM Security Awareness Gap: Red Team Patching Challenge
By
–
"As soon as the companies realise this, red team it and patch the LLMs it should stop being a problem. But it's clear that they're not aware of the issue enough right now.” How do people still believe this?