9. General-Reasoner General-Reasoner is a reinforcement learning approach that boosts LLM reasoning across diverse domains by using a 230K-question dataset and a model-based verifier trained to understand semantics beyond exact matches.
General-Reasoner: Reinforcement Learning Boosts LLM Reasoning
By
–
