MiniMax M3 just raised the bar—open weights, 1M-token context, and native multimodal from day one. Visit http://
futurepedia.io, the leading AI tools directory.
MULTIMODAL AI
-

MiniMax M3: Open weights, 1M tokens, native multimodal AI
By
–
-

Cosmos 3 merges reasoner and diffusion towers in one model
By
–
2/ What changed: old Cosmos split the work across separate models — one to understand a scene, one to generate video, one for controlled simulation.
Cosmos 3 fuses everything into a single Mixture-of-Transformers with two towers:
→ a reasoner (the VLM "brain")
→ a diffusion -
NVIDIA open-sources Cosmos 3, first open omnimodel for physical AI
By
–
1/ NVIDIA just open-sourced Cosmos 3 at GTC Taipei!
— Chubby♨️ (@kimmonismus) 1 juin 2026
It's the first fully open "omnimodel" for physical AI – one model that understands the real world, predicts what happens next, and generates the actions a robot should take.
Weights, code, datasets. All open. And this is… https://t.co/5Y8BbXUWqJ pic.twitter.com/3bOMlwO0B21/ NVIDIA just open-sourced Cosmos 3 at GTC Taipei! It's the first fully open "omnimodel" for physical AI – one model that understands the real world, predicts what happens next, and generates the actions a robot should take. Weights, code, datasets. All open. And this is
-
Ultra-realistic photo of a glamorous woman in a crowded stadium
By
–
Hecho con Seedance 2.0 en Pollo AI
— Nico (@nicos_ai) 1 juin 2026
Prompt:
"Foto ultrarrealista de retransmisión deportiva de una mujer glamurosa sentada entre el
público de un estadio de fútbol abarrotado durante un partido nocturno, con un top de satén sin
mangas de cuello alto en marrón oscuro y pendientes… pic.twitter.com/yR2fWosLGMMade with Seedance 2.0 on Pollo AI Prompt:
"Ultra-realistic sports broadcast photo of a glamorous woman sitting among the crowd of a packed football stadium during a night match, wearing a dark brown satin sleeveless turtleneck top and curls" -
NVIDIA accelerates bounding box detection 10x by removing mandatory token prediction
By
–
🚨 NVIDIA just pulled off something crazy: making bounding box detection 10x faster by ripping out the exact step the entire industry assumed was mandatory ↓
— Charly Wargnier (@DataChaz) 1 juin 2026
Every VLM grounding model treats boxes like sentences, predicting them token by token. It’s inherently slow.
Enter… pic.twitter.com/OE7fxZFF4VNVIDIA just pulled off something crazy: making bounding box detection 10x faster by ripping out the exact step the entire industry assumed was mandatory ↓ Every VLM grounding model treats boxes like sentences, predicting them token by token. It’s inherently slow. Enter
-
YOLOE-26: Three ways to detect objects – text, visual, prompt-free
By
–
YOLOE-26 turns object detection into three ways of saying "find this":
— Satya Mallick (@LearnOpenCV) 1 juin 2026
→ Text prompt (name it)
→ Visual prompt (show it)
→ Prompt-free (let the model decide)
Closed-set rigidity → open-vocabulary conversation.
Tutorial + benchmarks: https://t.co/od9zkfvMaX pic.twitter.com/t1bXsIl1HJYOLOE-26 turns object detection into three ways of saying "find this":
→ Text prompt (name it)
→ Visual prompt (show it)
→ Prompt-free (let the model decide)
Closed-set rigidity → open-vocabulary conversation.
Tutorial + benchmarks: https://
vist.ly/565gr -

New NVIDIA Cosmos 3 Model Provides Pretrained Foundation for Physical AI
By
–
Trained on billions of samples across modalities, the model provides developers with a powerful pretrained foundation for building physical AI systems with less data and lower training costs. Read more in our technical blog: https://
developer.nvidia.com/blog/develop-p
hysical-ai-reasoning-world-and-action-models-with-nvidia-cosmos-3/
… -
Nvidia AI generates F1 racing video from dashcam image
By
–
Image-to-video generation is just as impressive.
— NVIDIA AI (@NVIDIAAI) 1 juin 2026
Input image:
"Generate a 16:9 image from a dashcam view of a formula 1 racing event"
Video prompt:
"A high-speed racing event where a car navigates multiple winding turns"
🔊 Sound on – generated by Cosmos 3. pic.twitter.com/FxCWMVNDxhImage-to-video generation is just as impressive. Input image:
"Generate a 16:9 image from a dashcam view of a formula 1 racing event" prompt:
"A high-speed racing event where a car navigates multiple winding turns" Sound on – generated by Cosmos 3. -
Cosmos 3 excels at multimodal reasoning, simulation, and robot training
By
–
In addition to understanding and reasoning across modalities, Cosmos 3 excels at simulating physical environments, predicting future world states, and helping train robots to perform specific tasks.
— NVIDIA AI (@NVIDIAAI) 1 juin 2026
It can do subsecond vision reasoning, large scale synthetic data generation and… pic.twitter.com/UJ8dIEG5NDIn addition to understanding and reasoning across modalities, Cosmos 3 excels at simulating physical environments, predicting future world states, and helping train robots to perform specific tasks. It can do subsecond vision reasoning, large scale synthetic data generation and
-

NVIDIA AI ranks first on physical AI benchmarks among open models
By
–
Delivering leading results on physical AI benchmarks among open models, it ranks first across @ArtificialAnlys
, Physics-IQ, PAI-Bench and R-Bench for world generation accuracy, RoboLab and RoboArena for action policy and the VANTAGE-Bench and TAR leaderboards for vision