Oh I see. Yeah maybe depends on how extensive the preference tuning was. But in general it should be possible to achieve the same "flexibility" there as with non-reasoning LLMs.
Flexibility in Reasoning LLMs with Preference Tuning
By
–
By
–
Oh I see. Yeah maybe depends on how extensive the preference tuning was. But in general it should be possible to achieve the same "flexibility" there as with non-reasoning LLMs.