TPU 8i is co-designed with our Gemini research team to support low latency inference. Among the attributes that support this are large amounts of on-chip SRAM, enabling more computations to be done on chip without having to go to HBM for weights or KVCache state as often. The
COMPUTING
-
Operationalizing AI Deployment Across Defense Cloud and Edge
By
–
AI isn’t failing because of models; it’s failing at deployment. Join Spectro Cloud and @DominoDataLab to learn how defense teams operationalize AI across cloud, classified, and edge environments. April 28 | 1 PM ET Save your spot: https://
hubs.ly/Q04ctjBb0 Carahsoft -

TPU 8t: 3X FP4 Performance Boost Over Ironwood
By
–
First, let's talk about TPU 8t, which is designed for large-scale training and inference throughput. The pod size is increased slightly to 9600 chips, and provides ~3X the FP4 performance per pod vs. Ironwood (8t has 121 exaflops/pod vs. 42.5 exaflops/pod for Ironwood). In
-
Google announces eighth-generation TPU chips for agentic era
By
–
I had a good time discussing yesterday's Google TPU v8t and v8i announcement at Cloud Next with Amin Vahdat along with @AcquiredFM hosts @gilbert and @djrosent
. The blog post announcement has lots of details about these new chips: https://
blog.google/innovation-and
-ai/infrastructure-and-cloud/google-cloud/eighth-generation-tpu-agentic-era/
… Here's a thread of -
CPUs Critical Infrastructure for Agentic AI Performance
By
–
🔥 Hot take: CPUs don’t get enough credit in agentic AI.
— SambaNova (@SambaNovaAI) 23 avril 2026
They prep data, route requests, and coordinate with accelerators, while also handling everything outside the model like code execution, DB queries, and validation.
Without them, inference can’t keep up 🦾
Do you agree… pic.twitter.com/RzBjFUHACdHot take: CPUs don’t get enough credit in agentic AI. They prep data, route requests, and coordinate with accelerators, while also handling everything outside the model like code execution, DB queries, and validation. Without them, inference can’t keep up Do you agree
-

NVIDIA Advances Codex and GPT-5.5 Development at HQ
By
–
Codex + GPT-5.5 is moving fast here at NVIDIA HQ.
— NVIDIA AI (@NVIDIAAI) 23 avril 2026
We’ve even got our own Codex Lab for NVIDIANs to get started. https://t.co/PTqytJWp0f pic.twitter.com/zWnSnKfansCodex + GPT-5.5 is moving fast here at NVIDIA HQ. We’ve even got our own Codex Lab for NVIDIANs to get started.
-

CUDA Kernels and Custom Heuristic Routing Algorithms
By
–
"custom heuristic algorithms" was the cuda kernels yeah? wasnt exactly clear in the wording or if this part is hinting at some kind of cool routing work
-
kUPS GPU Optimization Achieves 49x Throughput Over RASPA
By
–
We’ve optimized kUPS specifically for GPU in collaboration with @nvidia , achieving up to 49× throughput over widely used software like RASPA for specific simulations.
-
Streamlining AI Workflows: GPU Efficiency and Integration Challenges
By
–
Real-world applications usually require a clunky, fragmented workflow: Chaining disparate, complex packages
Writing hundreds of lines of configuration
Poor GPU efficiency -
kUPS: Molecular Simulation Engine for AI Workflows
By
–
Today at @iclr_conf 2026, I was excited to announce kUPS: a molecular simulation engine built for the AI era, optimized for GPU in collaboration with NVIDIA. kUPS is a plug-and-play, Python-native toolkit designed to integrate seamlessly with modern ML workflows.
