6/ Fast Inference of Mixture-of-Experts – achieves efficient inference of Mixtral-8x7B models through offloading; designs a MoE-specific offloading strategy that enables running Mixtral-8x7B on desktop hardware and free-tier Google Colab instances.
Fast Inference of Mixture-of-Experts Models via Offloading
By
–
