To support efficient serving of Jamba-1.5-Large, we developed a novel quantization technique – ExpertsInt8. We quantize the MoE and MLP weights to INT8 in order to store them, and dequantize them back to BF16 before the actual computation. This technique is both very fast and
ExpertsInt8: Novel Quantization Technique for Jamba-1.5-Large Efficient Serving
By
–
