AI Dynamics

Global AI News Aggregator

About

RL Impact on Base Model Performance: Pre-training and Mid-training Interplay

There are competing views on whether RL can genuinely improve base model's performance (e.g., pass@128). The answer is both yes and no, largely depending on the interplay between pre-training, mid-training, and RL. We trained a few hundreds of GPT-2 scale LMs on synthetic GSM-like reasoning data from scratch. Here are what we found: 🧵

→ View original post on X — @jeande_d, 2025-12-09 20:20 UTC