What if LLM reinforcement learning could assign credit more accurately by thinking step by step? Researchers from Peking University and Microsoft Research Asia introduce GenAC: a generative critic that replaces one-shot value predictions with chain-of-thought reasoning before
GenAC: chain-of-thought generative critic for LLM RL
By
–
