AI Dynamics

Global AI News Aggregator

About

SlimQwen: Technical Research on Pruning and Distillation for MoE Models

“SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training” This new Qwen paper shows that pruning a pretrained MoE is much better than training the smaller MoE from scratch. All you need to do is prune depth, width, and experts, preserve some experts

→ View original post on X — @askalphaxiv