Among all publicly available and proprietary models, Jamba-1.5 models are the only ones with an effective length of 256K tokens, as evaluated by the RULER benchmark. 5/7
@ai21labs
-

Jamba Architecture: Mamba-1-Attention Hybrid Outperforms Mamba-2
By
–
We experimented with alternatives to our final Jamba architecture, including Mamba-2 (which was released a few months after the original Jamba). However, we found that in a hybrid architecture, the Mamba-1-Attention combination outperforms both Mamba-2-Attention and pure Mamba-2.
-

ExpertsInt8: Novel Quantization Technique for Jamba-1.5-Large Efficient Serving
By
–
To support efficient serving of Jamba-1.5-Large, we developed a novel quantization technique – ExpertsInt8. We quantize the MoE and MLP weights to INT8 in order to store them, and dequantize them back to BF16 before the actual computation. This technique is both very fast and
-

Jamba-1.5 Whitepaper Released: Hybrid SSM-Transformer Models
By
–
Jamba-1.5 whitepaper is out!
The whitepaper details the architecture, training schemes, novelties and in-depth evaluations of our new long context hybrid SSM-Transformer models – Jamba-1.5-Large and Jamba-1.5-Mini. Arxiv: https://
arxiv.org/abs/2408.12570 Here are some highlights and -

Jamba-1.5 Hybrid Architecture Delivers Superior Throughput and Latency Performance
By
–
The hybrid Jamba architecture enables Jamba-1.5 models to reach excellent throughput and latency, especially at long contexts. With the same hardware, Jamba-1.5 models are the fastest across the board (in the image: 2xA100 80GB GPUs for Mini, 8xA100 80GB GPUs for Large). 2/7
-
ExpertsInt8 Quantization Enables Jamba 1.5 Large on Single Node
By
–
To make sure Jamba 1.5 Large fits a single node, we developed ExpertsInt8, a quantization technique for MoE models. We are now able to fit the model with its entire 256K context window on a single node, without losing quality or having a complicated quantization process. [6/6]
-

Jamba 1.5 Models Excel on Arena Hard Benchmark
By
–
On Arena Hard benchmark, Jamba 1.5 Mini scores an excellent 46.1, surpassing even larger models. Jamba 1.5 Large scores 65.4, above Llama 3.1 70B and Llama 3.1 405B. [5/6]
-

Jamba 1.5 Mini Achieves Fastest Speed at 10K Contexts
By
–
Jamba’s speed also shows in OTPS, as seen in @ArtificialAnlys
. Jamba 1.5 Mini ranks the fastest @ 10K contexts [4/6] -

Jamba 1.5 Models Deliver 2.5X Faster Inference Performance
By
–
Both Jamba 1.5 models are faster than competitors of a similar size, with up to 2.5X faster inference on long contexts. [3/6]
-
Jamba 1.5 Models: Hybrid SSM-Transformer Architecture Innovation
By
–
The Jamba 1.5 models are based on our novel hybrid SSM-Transformer architecture, which combines the quality, speed and efficiency of both. Jamba 1.5 Mini has 12B active/52B total parameters, while Large is 94B active/398B total – the largest Mamba model ever made. [2/6]
