The models will acquire some amount of verbal reasoning just from the language modelling objective. But reinforcement learning training that specifically targets reasoning paths does give you something quite different.
Language Models Acquire Reasoning Through Reinforcement Learning Training
By
–