Perplexity runs on NVIDIA. Nice breakdown from the team on how they’re using the CUTLASS Python stack to optimize their models for inference
@nvidiaai
-
Building Vision AI Pipelines with Coding Agents and DeepStream
By
–
What if you could go from concept to a vision AI app without writing every line of code? ⚡
— NVIDIA AI (@NVIDIAAI) 7 mai 2026
NVIDIA DeepStream, combined with powerful coding agents like Claude Code and reusable Skills, can generate complete vision AI pipelines from simple natural language prompts, reducing… pic.twitter.com/Vh01ZseTokWhat if you could go from concept to a vision AI app without writing every line of code? NVIDIA DeepStream, combined with powerful coding agents like Claude Code and reusable Skills, can generate complete vision AI pipelines from simple natural language prompts, reducing
-

NVIDIA Research Introduces Guess-Verify-Refine Sparse-Attention Algorithm
By
–
What if every decode step gave the next one a head start? Meet Guess-Verify-Refine — a new hardware-aware sparse-attention algorithm from NVIDIA Research. Built for TensorRT LLM on Blackwell, it reuses temporal patterns across decode steps for: → 1.88x faster Top-K attention
-

NVIDIA AI announces TokenSpeed, a fast inference engine for agentic workloads
By
–

TokenSpeed is a brand new inference engine purpose built for speed-of-light agentic workloads. Read their blog to learn more about its advanced KV cache management, safe and efficient scheduler, and pluggable layered kernel system designed for multi-silicon support. Plus, it
-
Building Sub-Agents with NVIDIA Nemotron 3 Nano Omni Tutorial
By
–
How the Developer Community Builds Sub-Agents with NVIDIA Nemotron 3 Nano Omni | Nemotron Labs https://t.co/asHUSeUPgl
— NVIDIA AI (@NVIDIAAI) 5 mai 2026How the Developer Community Builds Sub-Agents with NVIDIA Nemotron 3 Nano Omni | Nemotron Labs
-

Scaling Agentic Workloads: 400+ Tokens/sec/User on Vera Rubin
By
–
What does it actually take to run agentic workloads at scale? Agents push token consumption, context length, and latency into extremely demanding regions. Extreme co-design on the Vera Rubin platform is built for these complex workloads, delivering 400+ tokens/sec/user on
-
NVIDIA Uses cuOpt Agentic Workflows to Optimize Supply Chains
By
–
Internally at NVIDIA, we use cuOpt based agentic workflows with agent skills to optimize our supply chains. Since it’s open source, you can too.
— NVIDIA AI (@NVIDIAAI) 4 mai 2026
With optimizations ready in minutes instead of weeks, the workflow uses multi-agent LangChain Deep agent orchestration and… pic.twitter.com/V5BJnMOzJbInternally at NVIDIA, we use cuOpt based agentic workflows with agent skills to optimize our supply chains. Since it’s open source, you can too. With optimizations ready in minutes instead of weeks, the workflow uses multi-agent LangChain Deep agent orchestration and
-
NVIDIA Megatron Core Adds Muon and Advanced Optimizers for LLM Training
By
–
Training Kimi K2 and Qwen3 30B-scale models efficiently requires more than standard data-parallel tricks. NVIDIA Megatron Core now provides end-to-end support for emerging higher-order optimizers like Muon, alongside research optimizers such as MOP and REKLS, to push training
-

Nemotron 3 Super Tops Open Source EnterpriseOps-Gym Leaderboard
By
–
Benchmarks should reflect real-world performance. That’s why we’re excited to share that Nemotron 3 Super has topped the open source category on the EnterpriseOps-Gym leaderboard. This agentic gauntlet evaluates performance across 1,150 tasks in fully interactive environments
-
OpenShell Open-Source Sandbox Makes AI Agents Safe for Enterprises
By
–
We created OpenShell to make AI agents safe for enterprises.
— NVIDIA AI (@NVIDIAAI) 1 mai 2026
Built in open source so any company can adopt and trust it, this secure sandbox controls what agents can access, share, and send.
Our CEO, Jensen, explains 👇 pic.twitter.com/7EiIsxr0CGWe created OpenShell to make AI agents safe for enterprises. Built in open source so any company can adopt and trust it, this secure sandbox controls what agents can access, share, and send. Our CEO, Jensen, explains
