You might find our new 3B reasonoing open model intersting. It's ranked above llma, Granite, Gemma, etc. on the tiny category. It has effective long context and excel in IF benchmarks.
@ai21labs
-
AI21 Labs Launches New 3B Reasoning Open Model Hybrid Architecture
By
–
Check out our new 3B reasonong open model – hybrid mamba nd transformer. It's ranked above Granite, Gemma, llma, etc. on the tiny category.
-
New 3B Jamba Model Ranked #1 on IFBench
By
–
You should try our new 3B Jamba. Ranked #1 on IFBench (tiny model category). It's an open [reasoning] model –
-
Jamba Reasoning 3B: Small Models Deliver Big AI Impact
By
–
Appreciate the feature, @IEEESpectrum
! Small models, big impact – that’s exactly what Jamba Reasoning 3B is all about. -
AI21 Labs Announces Jamba Reasoning 3B Release
By
–
Appreciate the shout-out, @AMULETAnalytics We’re thrilled to share Jamba Reasoning 3B with the world.
-
Jamba 1.5 Large Now Available for Local Inference Download
By
–
5/5 Available today for download & local inference on @huggingface
, @kaggle
, @lmstudio
, and llama.cpp. -
Jamba Achieves Superior Mobile Performance with Extended Context Lengths
By
–
4/5 The same efficiency gains apply on mobile. Running at 16K context lengths on an iPhone 16 Pro, Jamba outputs nearly 16 tokens/second, outpacing token outputs from Llama 3.2 3B, Qwen 3 1.7B, and Phi-4 Mini. Jamba is the only one that can handle up to 64K.
-

Jamba Reasoning 3B: Exceptional Performance on Extended Contexts
By
–
3/5 Where most other tiny models choke at context lengths above 8K, Jamba Reasoning 3B stays steady, with a consistent 30-40 tokens/second on an M3 MacBook Pro, regardless of context size. This is up to an order of magnitude faster than other on-device models.
-

Jamba Reasoning 3B Excels in Knowledge and Instruction-Following
By
–
2/5 Jamba Reasoning 3B excels in general knowledge benchmarks (MMLU-Pro, HLE) and instruction-following (IFBench), making it reliable and performant
-

Jamba Reasoning 3B: Hybrid SSM-Transformer Apache 2.0 Release
By
–
1/5 Releasing Jamba Reasoning 3B under Apache 2.0: Hybrid SSM-Transformer architecture that tops accuracy & speed across record context lengths. e.g. 3-5X faster than Llama 3.2 3B and Qwen3 4B at 32K tokens.