This explains why DeepSeek-R1 beats models 10x its size. It's not bigger. It's not trained on more data. It just thinks longer and verifies harder. 32B parameters thinking for 30 seconds > 405B parameters answering instantly. The scaling law just changed from "bigger" to
DeepSeek-R1 beats 10x larger models by thinking longer
By
–