This Meta paper just gave an empirical proof of why RL updates in LLMs are so sparse! They proposed this "Three-Gate Theory” & showed that pretrained models have highly structured optimization landscapes, so geometry-aware RL is much better than heuristic based SFT methods
Meta’s Three-Gate Theory Explains Sparse RL Updates in LLMs
By
–
