The most interesting part is that the base LFM-1B model isn't that strong in math (see results). Extensive SFT (~100B tokens) was enough to turn it into a strong reasoner. Further GRPO compressed the reasoning traces and even maintained performance.
SFT and GRPO Improve LFM-1B Math Reasoning Performance
By
–