We just shipped NVIDIA-Verified Agent Skills Skills make your agent more capable, but can also introduce vulnerabilities. Verified skills give you transparency into what a skill does, where it came from, what risks it carries, and whether it's been modified. Every verified
@nvidiaai
-
Release of Nemotron-Labs-Diffusion parallel generation language models
By
–
Most language models only generate one token at a time.
— NVIDIA AI (@NVIDIAAI) 19 mai 2026
We just released Nemotron-Labs-Diffusion, a family of diffusion language models that take a different approach, generating multiple tokens in parallel within a single model. Rather than committing to each token permanently,… pic.twitter.com/fTOBmQ8KaMMost language models only generate one token at a time. We just released Nemotron-Labs-Diffusion, a family of diffusion language models that take a different approach, generating multiple tokens in parallel within a single model. Rather than committing to each token permanently,
-

Technical Overview of Hybrid Linear Attention and Generation Pipeline
By
–
The architecture features Hybrid Linear Attention, Dual-Branch Camera Control, Two-Stage Generation Pipeline, and a Robust Annotation Pipeline. All of this combines for stronger action-following accuracy and higher throughput while maintaining visual quality.
-
Release of SANA-WM: Open Source World Model for Video Generation
By
–
One image + text + camera trajectory = controllable worlds. All on a single GPU.
— NVIDIA AI (@NVIDIAAI) 19 mai 2026
Our research team just released SANA-WM, a 2.6B open source world model natively trained for 60-second video generation with precise camera control. pic.twitter.com/oXHRCnCRdMOne image + text + camera trajectory = controllable worlds. All on a single GPU. Our research team just released SANA-WM, a 2.6B open source world model natively trained for 60-second video generation with precise camera control.
-
How to Authenticate Agent Skills Before Execution
By
–
How to Authenticate Agent Skills Before Execution | Nemotron Labs https://t.co/aSkgYYmNtB
— NVIDIA AI (@NVIDIAAI) 19 mai 2026How to Authenticate Agent Skills Before Execution | Nemotron Labs
-
OpenShell v0.0.41 Introduces Agent-Driven Policy Management
By
–
OpenShell v0.0.41 agent-driven policy management sandbox resource flags in the CLI custom CA support for OIDC TLS verification sandbox downloads with workspace-boundary checks bug fixes and stability improvements Policy and resource control, directly from the
-
Fastokens addresses tokenization bottlenecks in AI inference pipelines
By
–
Thanks for building with us @CrusoeAI As context windows explode, tokenization is becoming a major hidden bottleneck in inference pipelines. fastokens is open source, already integrated with Dynamo & @lmsysorg
, and designed for the next generation of 100K-token agent -
Optimizing Infrastructure for Agentic AI Inference
By
–
Delivering agentic inference at scale requires balancing efficiency across:
— NVIDIA AI (@NVIDIAAI) 13 mai 2026
1) Models and algorithms
2) Software
3) Compute
Our full-stack platform continuously optimizes for these inputs using extreme co-design across compute, networking, storage, and memory. Plus, software… pic.twitter.com/rzoF9wyF1NDelivering agentic inference at scale requires balancing efficiency across: 1) Models and algorithms
2) Software 3) Compute Our full-stack platform continuously optimizes for these inputs using extreme co-design across compute, networking, storage, and memory. Plus, software -

Hardening agentic stacks against reasoning and tool parsing drift
By
–
Most agentic stacks run into the same problems pretty quickly: reasoning and tool parsing drift across turns, KV cache reuse falls apart, or tools fire too late. We’ve been hardening Dynamo’s harness-facing path so @Claudeai Code, @OpenClaw
, and @openai Codex-style agent -

New ICML26 Paper on Sparse Transformer Kernels for NVIDIA GPUs
By
–
Great collab with @SakanaAILabs on an #ICML26 paper about sparse transformer kernels + formats optimized for modern NVIDIA GPU execution. • TwELL sparse packing
• Fused CUDA kernels
• 20%+ inference/training speedups at scale Paper + code below
