Grosse semaine en vue pour OpenAI aussi. Cette semaine s'annonce assez mouvementé. Les employés laissent entendre qu’un événement excitant se prépare cette semaine. – GPT-5 ?
– Un modèle open-weight ? Pendant ce temps, ChatGPT atteint les 700 millions d’utilisateurs
LLMS
-

OpenAI Teases Major Release as ChatGPT Hits 700 Million Users
By
–
-
Comprehensive Analysis of Model Architectures and Design Tradeoffs
By
–
We broke it all down: Key papers & model architectures Design tradeoffs: MoE, GQA, layer ordering Benchmarks across RULER, MMLU, ARC, HumanEval Open weights + distillation strategies
Read the full story here: -
Major AI Companies Release Mamba Architecture Models 2026
By
–
Since then, the space has exploded: @AI21Labs → Jamba & Jamba 1.5 @NVIDIAAI → MambaVision, Nemotron-H @MistralAI → Codestral Mamba @togethercompute → Mamba-Llama @IBMResearch → Bamba @TencentGlobal → Hunyuan TurboS @MSFTResearch →
-
Beyond Self-Attention: Scaling LLMs with Faster Inference and Alternative Architectures
By
–
Some pushed inference speed 5–10× faster.
Some scaled to 398B parameters.
Others rewired LLaMA-3 with Mamba layers—cutting latency without losing quality.
All of them moved beyond self-attention as the only tool for reasoning at scale. -
Mamba Paper: Revolutionary Foundation for Scalable LLM Architecture
By
–
It started with the original Mamba paper (Dec 2023) from @_albertgu & @tri_dao
:
→ Linear-time inference
→ Content-aware computation
→ Attention-free modeling That single paper cracked open a whole new path for scalable LLMs. -

Hybrid LLM Era: Beyond Transformers – Mamba to Bamba
By
–
Attention was never enough. The hybrid LLM era is here—and it’s moving fast. From Mamba to Jamba to Bamba, we mapped every major model that’s challenged the Transformer default in the past 18 months. A timeline of what’s changed and why it matters ↓
-
Iterating and Evaluating System Prompts for AI Models
By
–
Content strategy suggestion for @AnthropicAI or @OpenAI or any other AI Lab shipping systems with epic system prompts… I'd love to see a detailed write-up of how you iterate on, test and eval those prompts Would be a huge help for your customers who are building on your models
-

Google Teases Major Week Possibly Gemini 3.0 Launch
By
–

Le chef de produit IA de Google annonce un grosse semaine. Gemini 3.0 arrive ?! Je me souviens encore du monstre qu'était 2.5 à sa sortie. Cela risque de faire bouger Open AI
-
DeepSeek R1 Distills a 7B Model Surpassing GPT-4o on Reasoning
By
–
That DeepSeek R1 Distill example was surprising. Beating GPT-4o on reasoning with a 7B model isn’t small news.
-
LLM capabilities surge dramatically in past twelve months
By
–
A year ago, the vast majority of LLMs we had weren't even capable of telling us how many Rs STRAWBERRY has The leap in capabilities we've experienced since September 2024 has been enormous, and it seems we're going to close out these 12 months with a golden flourish.