"Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models" Instead of standard RL post-training collapsing an LLM toward one dominant answer, this paper shows you can train it to produce a set of plausible answers in a single pass. This is important because
RL Training for Distributional Reasoning in Language Models
By
–
