AI Dynamics

Global AI News Aggregator

About

SFT and GRPO Improve LFM-1B Math Reasoning Performance

The most interesting part is that the base LFM-1B model isn't that strong in math (see results). Extensive SFT (~100B tokens) was enough to turn it into a strong reasoner. Further GRPO compressed the reasoning traces and even maintained performance.

→ View original post on X — @maximelabonne