a 7M parameter model just beat DeepSeek R1’s 671 billion parameters on hard reasoning tasks. 45% vs 15.8%. trained in hours. fits in 28MB. runs on a single GPU. the trick is that it thinks twice. LLMs generate answers in one pass. make a mistake early, the whole thing
Small 7M Parameter Model Outperforms Large DeepSeek R1 on Reasoning
By
–
