0% markup + automatic model routing is the part people should pay attention to. The winner won’t be the app using the biggest model for everything.
It’ll be the one that knows when not to.
AI
-
Optimizing AI Application Performance via Intelligent Model Routing
By
–
-

Optimizing Throughput for Large MoE Models on GB 200 Hardware
By
–
GB 200s change how one does the prefill and decode disaggregation when serving large MoEs like Qwen. We’ve published details of our stack quantifying the throughput benefits compared to serving on Hoppers.
-

AGI ALPHA: A Scalable Substrate for Intelligence Organizations Manuscript Published
By
–
AGI ALPHA: A Scalable Substrate for Intelligence Organizations Vincent Boucher, President of http://
MONTREAL.AI and http://
QUEBEC.AI Paper: https://
github.com/MontrealAI/agi
alpha-first-real-loop/blob/main/docs/manuscript/AGI_ALPHA_Unified_Publication_Final.pdf
… #AGIAlpha #QuebecAI #SovereignAI -
NVIDIA GB200 Architecture Optimized for Large-Model Inference
By
–
This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization, custom kernels, and rack-scale NVLink turn GB200 into faster answers lower serving cost. Read the full paper here
-
Performance Benchmarks: NVIDIA H200 vs GB200 for AI Workloads
By
–
The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls from 730.1µs to 438.5µs. For decode, GB200 sustains much higher throughput at high token speeds.
-
Hardware Optimization for AI Model Prefill and Decode Phases
By
–
Prefill and decode stress hardware differently. Prefill is compute-bound, so Blackwell Tensor Cores, memory bandwidth, NVLink, and SHARP reductions help. Decode is latency/memory-bound, where GB200’s rack-scale NVLink domain opens up parallelism Hopper could not.
-

Research on Serving Qwen3 235B Models on NVIDIA GB200 Racks
By
–
We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up over Hopper for high-throughput inference on large MoE models, not just a training platform.
-
Debating the economic ROI of major AI labs
By
–
that’s actual speculative; so far OpenAI and Anthropic have lost money and most businesses have not gotten significant RoI
-
Gary Marcus’s Warnings on AI Generalization Proven True
By
–
“Marcus's repeated warnings about the "wall of generalization" since 1998 have once again been proven true.”
-
Google Gemini Omni Video Generation Model Leaks Ahead of I/O
By
–
🤯Google vient de leaker son prochain monstre.
— VISION IA (@vision_ia) 12 mai 2026
Gemini Omni : un modèle de génération vidéo qui vient d'apparaître chez certains utilisateurs Gemini, juste avant Google I/O (19-20 mai).
Les premiers résultats sont bluffants, un utilisateur a demandé "un professeur qui écrit une… pic.twitter.com/U86extcaNFGoogle vient de leaker son prochain monstre. Gemini Omni : un modèle de génération vidéo qui vient d'apparaître chez certains utilisateurs Gemini, juste avant Google I/O (19-20 mai). Les premiers résultats sont bluffants, un utilisateur a demandé "un professeur qui écrit une