new doc page dropped! rendering Agent Traces on the @huggingface Hub
AI
-
NVIDIA Launches Cosmos Coalition for Open World Models
By
–
7/ NVIDIA also launched the Cosmos Coalition – with Black Forest Labs, Runway, Skild AI, Agile Robots, LTX and Generalist – to push open world models forward together. The bet: world models are becoming the intelligence layer for robots and AVs, and NVIDIA wants the open
-
Cosmos 3 Nano and Super models on Hugging Face with datasets
By
–
6/ Two sizes, both live on Hugging Face right now: Cosmos 3 Nano (8B) — runs on a single workstation GPU for real-time robotics
Cosmos 3 Super (32B) — datacenter-grade, max quality Plus six open datasets and full post-training scripts on GitHub. -
Cosmos 3 tops open leaderboards for text, video, physics, and robotics
By
–
5/ And it's not a demo. Cosmos 3 tops the open leaderboards: #1 open model on Artificial Analysis for text→image AND image→video
#1 on Physics-IQ for physics accuracy — ahead of Sora 2
Leads PAI-Bench overall, ahead of Veo 3.1
#1 robot policy on RoboArena
Open weights beating -

Cosmos 3: Synthetic data generation for robotics, 60x faster, including rare edge cases.
By
–
4/ The real unlock is data. Robotics has always been bottlenecked by how little real-world training data exists. Cosmos 3 generates physically-accurate synthetic data up to 60x faster – including the rare edge cases you can't safely film: collisions, near-misses, accidents.
Eval -
Cosmos 3: native action generation for VLM, world model, robot
By
–
3/ Inputs and outputs span text, image, video, audio AND action.
— Chubby♨️ (@kimmonismus) 1 juin 2026
That last one is the big deal. Cosmos 3 was trained natively to generate actions, so the same checkpoint can run as a vision-language model, a video world model, or a robot policy. No multi-model orchestration. pic.twitter.com/PoLsK33ytZ3/ Inputs and outputs span text, image, video, audio AND action. That last one is the big deal. Cosmos 3 was trained natively to generate actions, so the same checkpoint can run as a vision-language model, a video world model, or a robot policy. No multi-model orchestration.
-

MiniMax M3: Open weights, 1M tokens, native multimodal AI
By
–
MiniMax M3 just raised the bar—open weights, 1M-token context, and native multimodal from day one. Visit http://
futurepedia.io, the leading AI tools directory. -

Cosmos 3 merges reasoner and diffusion towers in one model
By
–
2/ What changed: old Cosmos split the work across separate models — one to understand a scene, one to generate video, one for controlled simulation.
Cosmos 3 fuses everything into a single Mixture-of-Transformers with two towers:
→ a reasoner (the VLM "brain")
→ a diffusion -
NVIDIA open-sources Cosmos 3, first open omnimodel for physical AI
By
–
1/ NVIDIA just open-sourced Cosmos 3 at GTC Taipei!
— Chubby♨️ (@kimmonismus) 1 juin 2026
It's the first fully open "omnimodel" for physical AI – one model that understands the real world, predicts what happens next, and generates the actions a robot should take.
Weights, code, datasets. All open. And this is… https://t.co/5Y8BbXUWqJ pic.twitter.com/3bOMlwO0B21/ NVIDIA just open-sourced Cosmos 3 at GTC Taipei! It's the first fully open "omnimodel" for physical AI – one model that understands the real world, predicts what happens next, and generates the actions a robot should take. Weights, code, datasets. All open. And this is
-

GrepSeek: Training Search Agents for Direct Corpus Interaction
By
–
GrepSeek Training Search Agents for Direct Corpus Interaction
