Maybe a container issue. I know Nvidia recommends their docker containers, but I am usually just running it on the Spark directly.
HARDWARE
-

Google Researcher: AI Crisis Lies in Inference, Not Training
By
–
BREAKING: A Google researcher and a Turing Award winner just published a paper arguing that the real AI crisis isn’t training. It’s inference. And the hardware stack we rely on today was never built for it. The paper, by Xiaoyu Ma and David Patterson, was accepted by IEEE
-
SN50 RDU Chip Promises 5X Faster Agentic AI Inference
By
–
SN50 RDU — purpose-built for agentic inference. Max speed of up to 5X faster; run agentic AI at a 3X lower cost than GPUs, unlocking cloud-scale inference economics. Learn more in our blog
-

Sunday Robotics Raises $165M Series B, Becomes Unicorn
By
–
BREAKING: @sundayrobotics has just raised a MASSIVE $165M Series B at a $1.15 billion valuation! It has happened! @Coatue led the round, with @BainCapVC, TigerGlobal, @benchmark, and FidelityInv also participating. Thomas Laffont joins the board! This has been a long time coming and must be a great feeling for those working at @sundayrobotics. The company emerged from stealth just four months ago and is now already a unicorn. @tonyzzhao and @chichengcc dropped out of Stanford PhDs to build what could become the first real household robot. Being able to deploy their robot Memo into real homes later this year through their Beta program could be a huge inflection point for the entire home robotics space. Credit: sunday.ai Congrats to @tonyzzhao, @chichengcc and the entire team!! 👋
→ View original post on X — @ken_goldberg, 2026-03-12 17:43 UTC
-

AI Inference for Real-World Constraints and Safety-Critical Applications
By
–
Axelera AI x Meridian Data Labs Defense installations. Remote industrial sites. Aircraft and engine MROs. Safety-critical metrology inspections. X-ray baggage inspections. These are environments where AI inference must operate under real-world constraints, not in datacenter
-

Scalable MoE Training Efficiency with Megatron Core
By
–
“Scalable Training of Mixture-of-Experts Models with Megatron Core” This NVIDIA MoE report walks through the hard part of MoE training. The key is not to add more parameters, but keeping sparse models efficient when only a small part of the model runs for each token. For
-

CUDA Agent: Large-Scale Agentic RL for Kernel Generation
By
–
CUDA Agent | Large-Scale Agentic RL for CUDA Kernel Generation https://
buff.ly/13zVaiQ
#AI #MachineLearning #DeepLearning #LLMs #DataScience -
Token Generation Speed Comparison: OSS and Nemotron Models Performance
By
–
Yeah. Not quite as bad for me, but:
gpt-oss-120b: 42.22 tokens/sec
nemotron-3-super: 20.43 tokens/sec on the DGX Spark. But this is ollama and might be an implementation issue. I have yet to try the Nvidia-optimized llama.cpp version (they had one for Nemotron Nano back then) -

llmfit: Auto-detect hardware and rank 206 models by VRAM compatibility
By
–
Stop guessing which models fit in your VRAM! llmfit is a CLI tool that auto-detects your hardware and ranks 206 models by what actually runs on your system. You download a 70B model and hope it fits. Or you estimate memory requirements across quantization levels and still end
-
India Launches National Compute Pool with Seven Cloud Providers
By
–
In 18 months, seven empanelled cloud providers. @NVIDIA H100s, @AMD MI300s, @Google Trillium TPUs. Distributed across tier-3 data centres. All under one national compute pool. (3/8) @AshwiniVaishnaw @jitinprasada @PIB_India @SecretaryMEITY @abhish18 @kavitabha @GoI_MeitY