5/ WARM – introduces weighted averaged rewards models (WARM) that involve fine-tuning multiple rewards models and then averaging them in the weight space; improves efficiency while improving the quality and alignment of LLM predictions.
WARM: Weighted Averaged Rewards Models Improve LLM Alignment
By
–