Some pushed inference speed 5–10× faster.
Some scaled to 398B parameters.
Others rewired LLaMA-3 with Mamba layers—cutting latency without losing quality.
All of them moved beyond self-attention as the only tool for reasoning at scale.
@ai21labs
-
Beyond Self-Attention: Scaling LLMs with Faster Inference and Alternative Architectures
By
–
-
Major AI Companies Release Mamba Architecture Models 2026
By
–
Since then, the space has exploded: @AI21Labs → Jamba & Jamba 1.5 @NVIDIAAI → MambaVision, Nemotron-H @MistralAI → Codestral Mamba @togethercompute → Mamba-Llama @IBMResearch → Bamba @TencentGlobal → Hunyuan TurboS @MSFTResearch →
-
Mamba Paper: Revolutionary Foundation for Scalable LLM Architecture
By
–
It started with the original Mamba paper (Dec 2023) from @_albertgu & @tri_dao
:
→ Linear-time inference
→ Content-aware computation
→ Attention-free modeling That single paper cracked open a whole new path for scalable LLMs. -

Hybrid LLM Era: Beyond Transformers – Mamba to Bamba
By
–
Attention was never enough. The hybrid LLM era is here—and it’s moving fast. From Mamba to Jamba to Bamba, we mapped every major model that’s challenged the Transformer default in the past 18 months. A timeline of what’s changed and why it matters ↓
-
Maestro integrates planning technology to improve reasoning
By
–
Exactly! That’s why we built planning technology into Maestro – to eliminate the reasoning point of failure.
-
Auditable Plans for Agent Transparency and Accountability
By
–
These steps should be part of a visible, auditable plan, an “auditable artifact”, as Karpathy says, that helps users and builders understand exactly how the agent tackled the work. 6/6
-
Rethinking AI System Architecture for User Requirements
By
–
It’s time to rethink how we architect AI systems. They should guarantee fulfillment of explicit user requirements, know how to break down work into the right steps, use the optimal tools to execute each step, and validate the output of each step. 5/6
-
AI Control: Putting Artificial Intelligence on a Leash
By
–
The solution? In the words of Karpathy, we need to “put AI on a leash.” 4/6
-
AI Challenges: Serious Work vs Vibe Coding
By
–
This is a good example of the challenges that arise when you want to use AI for serious work vs “vibe coding.” 3/6
-
Karpathy warns against excessive enthusiasm for autonomous AI agents
By
–
Andrej Karpathy recently warned we're getting "way too excited" about autonomous AI agents. “If I’m just vibe coding AI is great, but if I’m trying to really get work done, it’s not so great to have overreactive agents.” 2/6