So you want to see if your DP synthetic data method is actually any good. What makes a good benchmark? 1. Zero-shot performance should be low: the method should measure learning from the actual data;
2. Training on real data should work: learning should be possible; and… 2/n
GENERATIVE AI
-
Benchmarking DP synthetic data: zero-shot and real data training
By
–
-

ContinuousBench: Hard Leakage-Proof DP Synthetic Text Benchmark
By
–
Does DP synth text transfer useful knowledge or just superficial style mimicking? Existing benchmarks: saturated Introducing ContinuousBench: a hard (curr methods fail at ε=100! ) & leakage-proof benchmark for DP synth text! Followup to our #ICML2024 best paper 1/n
-

Claude Code deletes session traces after a month
By
–
i was today years old when i learned that claude code deletes your session traces after a month
-
Synthetic Data as Key to Making Robots Truly Usable
By
–
synthetic data is the key to make robots really usable
-
Cosmos 3 Nano and Super models on Hugging Face with datasets
By
–
6/ Two sizes, both live on Hugging Face right now: Cosmos 3 Nano (8B) — runs on a single workstation GPU for real-time robotics
Cosmos 3 Super (32B) — datacenter-grade, max quality Plus six open datasets and full post-training scripts on GitHub. -
Cosmos 3 tops open leaderboards for text, video, physics, and robotics
By
–
5/ And it's not a demo. Cosmos 3 tops the open leaderboards: #1 open model on Artificial Analysis for text→image AND image→video
#1 on Physics-IQ for physics accuracy — ahead of Sora 2
Leads PAI-Bench overall, ahead of Veo 3.1
#1 robot policy on RoboArena
Open weights beating -
Cosmos 3: native action generation for VLM, world model, robot
By
–
3/ Inputs and outputs span text, image, video, audio AND action.
— Chubby♨️ (@kimmonismus) 1 juin 2026
That last one is the big deal. Cosmos 3 was trained natively to generate actions, so the same checkpoint can run as a vision-language model, a video world model, or a robot policy. No multi-model orchestration. pic.twitter.com/PoLsK33ytZ3/ Inputs and outputs span text, image, video, audio AND action. That last one is the big deal. Cosmos 3 was trained natively to generate actions, so the same checkpoint can run as a vision-language model, a video world model, or a robot policy. No multi-model orchestration.
-

MiniMax M3: Open weights, 1M tokens, native multimodal AI
By
–
MiniMax M3 just raised the bar—open weights, 1M-token context, and native multimodal from day one. Visit http://
futurepedia.io, the leading AI tools directory. -

Cosmos 3 merges reasoner and diffusion towers in one model
By
–
2/ What changed: old Cosmos split the work across separate models — one to understand a scene, one to generate video, one for controlled simulation.
Cosmos 3 fuses everything into a single Mixture-of-Transformers with two towers:
→ a reasoner (the VLM "brain")
→ a diffusion -
NVIDIA open-sources Cosmos 3, first open omnimodel for physical AI
By
–
1/ NVIDIA just open-sourced Cosmos 3 at GTC Taipei!
— Chubby♨️ (@kimmonismus) 1 juin 2026
It's the first fully open "omnimodel" for physical AI – one model that understands the real world, predicts what happens next, and generates the actions a robot should take.
Weights, code, datasets. All open. And this is… https://t.co/5Y8BbXUWqJ pic.twitter.com/3bOMlwO0B21/ NVIDIA just open-sourced Cosmos 3 at GTC Taipei! It's the first fully open "omnimodel" for physical AI – one model that understands the real world, predicts what happens next, and generates the actions a robot should take. Weights, code, datasets. All open. And this is