Instead of evaluating each token, one at a time, RLFT grades entire answers—perfect for scenarios like OpenAI’s o1, where we can’t access intermediate tokens before the “end_of_thought.”
RLFT Evaluates Complete Answers Instead of Individual Tokens
By
–