To make sure Jamba 1.5 Large fits a single node, we developed ExpertsInt8, a quantization technique for MoE models. We are now able to fit the model with its entire 256K context window on a single node, without losing quality or having a complicated quantization process. [6/6]
ExpertsInt8 Quantization Enables Jamba 1.5 Large on Single Node
By
–