Why this matters: human written data & human reinforcement may set the upper bounds on what LLMs can do. AlphaGo was able to beat humans because it trained by playing against itself. This suggests that LLMs may be able to do similar self-play, offering a path to rapid improvement
Self-Play Training Could Overcome Human Data Limits for LLMs
By
–