8/ DeepSeekMath – continues pretraining a code base model with 120B math-related tokens; introduces GRPO (a variant to PPO) to enhance mathematical reasoning and reduce training resources via a memory usage optimization scheme.
DeepSeekMath Enhances Mathematical Reasoning with GRPO Optimization
By
–
