
At NVIDIA Trying to get my hands on some more GPUs

By
–
Google just turned Android into an AI delivery vehicle at its #TheAndroidShow event. New Gemini-native laptops. The phone OS is becoming a system-wide intelligence layer. The cursor is now an agent. It was a big day of releases, and I/O isn't even here yet. The most notable
By
–
Demos in the local AI space between heterogenous devices are misleading -Stuff like mixing MacBooks / Mac Studios with GPUs / DGX Sparks I haven’t seen a single mature software stack in that space yet To talk about it with high certainty is quite dishonest IMHO My 2 cents

By
–

Long energy, long nuclear, long orbital compute. We need more raw resources, electricity, infrastructure, and, critically, efficiency, throughout the entire supply chain of AI. The aim is to scale non-biological Intelligence so it can be distributed like a Utility. The signal

By
–
These are the Intelligence Factories: over 70% of these AI compute clusters are being built in the United States.
By
–
We're heading to Computex 2026 in Taipei at the beginning of next month, and we've got a lot to show you. Put our booth on your itinerary now, so you won't miss out on: Europa running a VLM to help find lost people across video streams A robot arm running object detection

By
–
GB 200s change how one does the prefill and decode disaggregation when serving large MoEs like Qwen. We’ve published details of our stack quantifying the throughput benefits compared to serving on Hoppers.
By
–
This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization, custom kernels, and rack-scale NVLink turn GB200 into faster answers lower serving cost. Read the full paper here
By
–
The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls from 730.1µs to 438.5µs. For decode, GB200 sustains much higher throughput at high token speeds.