A multimodal model reads a chart, misreads one number, then writes an object that is not in the image — and still arrives at the correct final answer. Result-scored RL rewards it anyway. Conversely, a wrong final answer does not mean every earlier observation was wrong: valid
Multimodal model misreads chart yet RL rewards correct answer
By
–
