It means combining different hardware architectures to split shards / run models across all of them at the same time
HARDWARE
-
GB300 NVL72 delivers 1.6x performance on DeepSeekV3 pretraining
By
–
Correction: GB300 NVL72 delivers 1.6x performance on DeepSeekV3 pretraining at 512 GPU scale. (Ref: results 6.0-0022 and 6.0-0101)
-
Models Will Disrupt Workflows; Infrastructure and Hardware Are Permanent Moats
By
–
Models will eat into Workflows Infra and Hardware are forever
always needed, always evolving Everything else has no real moat -

IneffableLabs builds super-learner with NVIDIA Vera Rubin NVL72
By
–
@IneffableLabs builds a super-learner that learns from experience, requiring enormous compute scale. They chose NVIDIA Vera Rubin NVL72, one of the largest clusters on @GoogleCloud, to power their mission to achieve superintelligence.
-
NVIDIA Blackwell platform dominates MLPerf Training 6.0 benchmarks
By
–
The NVIDIA Blackwell platform just swept MLPerf Training 6.0, delivering fastest performance and largest scale.
— NVIDIA (@nvidia) 16 juin 2026
Beyond the benchmarks, capabilities like the Reliability, Availability, and Serviceability Engine and NVIDIA Resiliency Extension deliver fewer interruptions and… pic.twitter.com/QH77j4UA8nThe NVIDIA Blackwell platform just swept MLPerf Training 6.0, delivering fastest performance and largest scale. Beyond the benchmarks, capabilities like the Reliability, Availability, and Serviceability Engine and NVIDIA Resiliency Extension deliver fewer interruptions and
-
Colossus 2 not fully used with Grok
By
–
To my knowledge, Colossus 2 is also not fully used with Grok.
-
Small specialized models are the future; buying a GPU was right
By
–
Small and specialized models are the future Buy a GPU was right
-
AI Inference: System-Level Time Problem in Post-Moore Era
By
–
My takeaway: AI inference is no longer just a model problem. It's a system-level time problem. For enterprise leaders, AI performance and AI economics are becoming inseparable. Learn more about Tau Scaling and what it means for the post-Moore era: https://
chinaxiv.org/abs/202605.002
24?locale=en
… What -
Why AI inference prioritises low latency over raw compute
By
–
Why does this matter for AI inference specifically? Training = throughput problem. Inference = latency problem. When a user talks to an AI assistant, tokens have to return fast. Latency, memory access, bandwidth, and interconnect all matter, not just raw compute. In large AI
-
Huawei’s Tau Scaling Law reframes AI performance bottleneck
By
–
Most AI teams are optimizing the model.
— Ronald van Loon (@Ronald_vanLoon) 16 juin 2026
But the real bottleneck in inference is underneath it.
Huawei's Tau Scaling Law (Her's Law) was just introduced at IEEE ISCAS in Shanghai.
It reframes how we think about AI performance entirely.
Here's the breakdown…#HuaweiPartner… pic.twitter.com/YihGFx75cyMost AI teams are optimizing the model. But the real bottleneck in inference is underneath it. Huawei's Tau Scaling Law (Her's Law) was just introduced at IEEE ISCAS in Shanghai. It reframes how we think about AI performance entirely. Here's the breakdown… #HuaweiPartner