AI is hitting a new phase and infrastructure is at the center of it. Don't forget to join @RodrigoLiang at @WebSummit Vancouver as he breaks down what’s changing and what comes next. May 13 | 12:20–12:40 PM https://
vancouver.websummit.com/sessions/van26
/d28d9215-dcdd-4b20-94e3-03d56a46061c/building-ai-s-400b-backbone/
…
COMPUTING
-

AI Infrastructure and Future Trends Discussed at Web Summit
By
–
-
AI Coder Builds OS with Claude Code: Masterclass, Guide, and Free Repo
By
–
This guy literally built an ENTIRE operating system with Claude Code 🤯
— Charly Wargnier (@DataChaz) 12 mai 2026
now he is showing you exactly how.
Nate just dropped the ultimate AI coding starter pack:
→ 2+ hour masterclass
→ Complete framework guide
→ Free GitHub repo to start building
bookmark this 👇 https://t.co/DwToTNzjwH pic.twitter.com/Atk3Nd4croThis guy literally built an ENTIRE operating system with Claude Code now he is showing you exactly how. Nate just dropped the ultimate AI coding starter pack: → 2+ hour masterclass
→ Complete framework guide
→ Free GitHub repo to start building bookmark this -

Optimizing Throughput for Large MoE Models on GB 200 Hardware
By
–
GB 200s change how one does the prefill and decode disaggregation when serving large MoEs like Qwen. We’ve published details of our stack quantifying the throughput benefits compared to serving on Hoppers.
-
NVIDIA GB200 Architecture Optimized for Large-Model Inference
By
–
This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization, custom kernels, and rack-scale NVLink turn GB200 into faster answers lower serving cost. Read the full paper here
-
Performance Benchmarks: NVIDIA H200 vs GB200 for AI Workloads
By
–
The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls from 730.1µs to 438.5µs. For decode, GB200 sustains much higher throughput at high token speeds.
-

Research on Serving Qwen3 235B Models on NVIDIA GB200 Racks
By
–
We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up over Hopper for high-throughput inference on large MoE models, not just a training platform.
-

Multi-Agent Synergy for Scaling Test-Time Compute
By
–
TMAS Scaling Test-Time Compute via Multi-Agent Synergy
-

Chip Complexity: The Era of Specialized Processors
By
–
Computational complexity has reached a level where a single general-purpose processor can no longer efficiently handle all AI tasks. To overcome bottlenecks, the industry has turned to specialized chips, each optimized for a specific type of workload. Microblog @antgrasso
-
Ilya Sutskever on the extreme compute costs of AI development
By
–
And right before that, when @ilyasut explained the cost of all the compute: “You look at this number of dollars and you say ‘gosh that’s a lot of dollars.’”
-

Human-AI Bandwidth Bottleneck in Modern ML Accelerators
By
–
In modern ML accelerators, FLOPS have absolutely exploded. Often though, the bottleneck is not FLOPS but memory bandwidth. Similarly, model intelligence has exploded, causing the bottleneck to be humanAI bandwidth. At Thinky, we think that it’s important to solve this. 1/4 x.com/thinkymachines…