Yeah I think it’s very similar! One difference I think is that in the RL they’re explicitly exposed to lots of deep math / coding problems with detailed solutions that are presumably not in the regular base model training data. And the reasoning models can expend more thinking
Reasoning Models Training with Deep Math and Coding Problems
By
–