Their hit was fine-tuning LLMs (“RL”), and this is still that. (Even CoT came from Google.)
Fine-tuning LLMs: The Continued Focus on Reinforcement Learning
By
–
By
–
Their hit was fine-tuning LLMs (“RL”), and this is still that. (Even CoT came from Google.)