i guess you could call it a dishwashing robot…
HARDWARE
-

Chiplet Architecture Replaces Transistor Scaling for Heat Management
By
–
Transistor scaling is hitting thermal and cost ceilings, pushing chip design toward modular chiplets connected by high-speed links. With shared substrates and tailored process nodes, we manage heat and sustain performance without relying on miniaturization. Microblog @antgrasso
-
Hyperscalers Leverage Nuclear Energy for AI Infrastructure
By
–
Godspeed to those engineers at RR, they have their work cut out for them. PS, in the U.S., the hyperscalers are treating energy as part of their tech stack. Microsoft is restarting 3-Mile Island, and Google / Meta et all are investing heavily in the next generation of SMR.
-
AI race hinges on scaling power and chip output
By
–
The AI race will come down to scaling power and chip output
-
Qwen 3.5 27B Matches Sonnet 4.5 on Single RTX 5090
By
–
That model, Qwen 3.5 27B Dense is equal to Sonnet 4.5 Runs great on a single RTX 5090 w/ full context But we are not anywhere near Opus 4.5 even with Qwen 3.5 397B MoE
-
Full Model GPU Offload Required to Avoid Performance Bottlenecks
By
–
Partial offloads wouldn’t work because this will create a severe bottleneck That’s why the full model has to be offloaded to the GPU to see actual performance gains
-
GPU Offloading Offers 1800Gbps vs DGX Spark’s 273Gbps
By
–
Yeah that is definitely faster because you’re fully offloading to the GPU which runs at ~1800Gbps in comparison to DGX Spark’s 273Gbps
-

New Website to Benchmark Hardware Tokens per Second Performance
By
–
Update: @AlpinDale and I agreed to collab on a website that does more accurate calculations for your hardware tokens/sec performance Expect more on this after I am done with GTC this week
-
Qwen 3.5 MoE Model Recommendations for Spark and Mac Studio
By
–
For a single Spark my first recommendation would be Qwen 3.5 122B MoE (Int4) For the Mac Studio I would recommend the 397B in 4-bit (and I think @ivanfioravanti would agree with me here)
-

Use MoE Models on Unified Memory Hardware Like DGX Spark
By
–
As I have mentioned before, stop trying to get Dense models running on the DGX Spark/Mac Studios Unified Memory is best fit for MoEs because you only make each token go through a small subset of the numbers of parameters in the model Optimize for your hardware
