7). Self-Play Preference Optimization – proposes a self-play-based method for aligning language models; this optimization procedure treats the problem as a constant-sum two-player game to identify the Nash equilibrium policy.
Self-Play Preference Optimization for Language Model Alignment
By
–
