AI Dynamics

Global AI News Aggregator

About

ExpertsInt8: Novel Quantization Technique for Jamba-1.5-Large Efficient Serving

To support efficient serving of Jamba-1.5-Large, we developed a novel quantization technique – ExpertsInt8. We quantize the MoE and MLP weights to INT8 in order to store them, and dequantize them back to BF16 before the actual computation. This technique is both very fast and

→ View original post on X — @ai21labs