They are! But both are mitigated by RL-based tuning and have similar costs in diversity and creativity. To tune for correctness is to optimize for tokens that, in practice, lead to true output, regardless of how well those choices represent real text. Consider whether an LLM
RL tuning mitigates issues with diversity and creativity cost
By
–