Interestingly, Hunyuan-Large follows a traditional post-training pipeline with SFT and DPO. SFT data is sourced from public sources and evolved to increase complexity Data quality is ensured through 3 stages: rules → 70B critique model → human review If I understood
Hunyuan-Large Post-Training Pipeline: SFT and DPO Strategy
By
–
