Announced at #GoogleCloudNext: NVIDIA and @GoogleCloud are bringing #agenticAI on-prem with support for Google Gemini models—powered by NVIDIA Blackwell and Confidential Computing for top-tier performance and security. https://
nvda.ws/4loZnhh Hear more from our CEO Jensen Huang.
HARDWARE
-

NVIDIA Google Cloud Deploy Agentic AI On-Premises Gemini
By
–
-

Llama 4 Achieves Record 2611 Tokens Per Second on Cerebras
By
–
Llama 4 is now live on Cerebras!
– We broke the perf chart again at 2,611 tokens/s
– 19x faster than the leading GPU cloud
– Only API in the world with <1s total response time
Try now: https://
inference.cerebras.ai -
Google Releases Gemini Agentic AI on Distributed Cloud
By
–
Announced at #GoogleCloudNext Enterprises can now unlock the full potential of #AgenticAI with the latest Gemini models now available on Google Distributed Cloud with NVIDIA Confidential Computing on NVIDIA Blackwell infrastructure. Learn more
-

Llama 4 Maverick fastest inference 655 tokens per second
By
–
I feel the need, the need for speed! Llama 4 Maverick from @AIatMeta is now available on SambaNova Cloud & it's the fastest inference verified by @ArtificialAnlys at 655 t/s. $0.50 / million input tokens & $2.00 per million output tokens. On SambaNova Cloud
-

Metis M.2 Powers Offline AI Chatbot on Arduino Portenta X8
By
–
Metis M.2 Meets Arduino Portenta X8 Check out the power of @AxeleraAI
’s Metis PCIe on the @arduino Portenta X8 at #ArduinoDay! Fully offline AI chatbot demo—no cloud needed. Now in Early Access: https://
eu1.hubs.ly/H0j8JM40 #EdgeAI #AIinference #Metis #Arduino #OfflineAI -

Understanding GPU Architecture Fundamentals for AI Systems
By
–
Fundamentals of GPU Architecture https://
buff.ly/Zc6p5ks
#AI #MachineLearning #DeepLearning #LLMs #DataScience -
Google Introduces Ironwood TPU for Inference Era
By
–
Introducing Ironwood, the first TPU built for the age of inference, and the timing could not be better : ) – Ironwood perf/watt is 2x relative to Trillium, 6th gen TPU
– Ironwood offers 192 GB per chip, 6x that of Trillium
– 4.5x faster data access -
Exponential AI Demand Requires Investment Across Full Stack
By
–
The demand for AI is on an exponential, we need to continue investing at every layer of the stack if we are going to be able to service the demand for AI compute efficiently and at humanity scale. TPU's are so wonderful : )
-

Open-Source Quadruped Robot: From CGI to Garage Build
By
–
wow this is soooo fun!
— Thomas Wolf (@Thom_Wolf) 9 avril 2025
CGI for now but really not far from current robot capabilities –though battery life is a challenge with 150 lb sitting on it–
now we need this as open-source hardware + software so I can build one in my garage pic.twitter.com/QzOMZkBLf3wow this is soooo fun! CGI for now but really not far from current robot capabilities –though battery life is a challenge with 150 lb sitting on it– now we need this as open-source hardware + software so I can build one in my garage
-
China Trade Impact: iPhone Prices Expected to Double Globally
By
–
You'll be fine? iPhones will 2x in prices, so will much likely 95% of all material good on your shelves. Since China basically provides the world with most made goods.