Cosmos 3 ties everything together. Previous releases separated world generation, physical understanding, and controlled scene generation. Cosmos 3’s MoT architecture unifies these capabilities by pairing an autoregressive reasoner tower with a diffusion-based generator tower.
MULTIMODAL AI
-
NVIDIA AI launches Cosmos 3, first fully open omnimodel for Physical AI
By
–
Introducing Cosmos 3: Our latest frontier model for Physical AI
— NVIDIA AI (@NVIDIAAI) 1 juin 2026
Cosmos 3 is the world’s first fully open omnimodel with native vision reasoning, world and action generation.
Today we’re releasing Super (32B) and Nano (8B) variants. pic.twitter.com/6UfkSA7kzQIntroducing Cosmos 3: Our latest frontier model for Physical AI Cosmos 3 is the world’s first fully open omnimodel with native vision reasoning, world and action generation. Today we’re releasing Super (32B) and Nano (8B) variants.
-
AI Talk to Art to Robot Build: Future Awesome
By
–
talk to AI -> create AI art -> robot builds it. Future gonna be so awesome
-
Photo tourism papers applied to personal iCloud/Google Photos
By
–
yup! photo tourism esque papers but applied to your own icloud/google photos
-
Guild: unifies Claude Code, Cursor, and Codex on one machine
By
–
Claude Code, Cursor, and Codex can now work as one team.
— AlphaSignal AI (@AlphaSignalAI) 31 mai 2026
Multi-agent coding setups keep losing context across tools and sessions.
Guild is a single Go binary that fixes this.
It runs a local MCP server backed by embedded SQLite.
Nothing leaves your machine, no cloud, no… pic.twitter.com/4JI1V4InrfClaude Code, Cursor, and Codex can now work as one team. Multi-agent coding setups keep losing context across tools and sessions. Guild is a single Go binary that fixes this. It runs a local MCP server backed by embedded SQLite. Nothing leaves your machine, no cloud, no
-

Can AI See and Hear Simultaneously?
By
–
Can AI truly see and hear together? Researchers from NUS, Oxford, University of Toronto, and Microsoft Research present a comprehensive survey on Audio-Visual Intelligence. They unify the fragmented field of large foundation models that combine sound and vision—from speech.
-

Anthropic plans consumer and bioscience expansion with new features
By
–
Anthropic is planning to further expand into the consumer and bioscience sectors. The biggest things to watch for – Conway agent
– Orbit assistant
– Knowledge-based memory
– Multilingual Voice Mode
– Operon for bioscience researchers and more! Which one do you think will -
Half-asleep post: GPT Image 2 fails to follow “Minimalist flat-lay”
By
–
My bad, it was posted while I was half asleep TBH, from my pov GPT Image 2 didn't follow the most important part at all "Minimalist flat-lay"
-
Using red grid lines to help models navigate screenshots
By
–
For example, models could not really “see” well. So one thing I used to do before injecting a computer screenshot into the model was overlay red grid lines on top, so the model could see the sections better and know where to click. That’s why this was able to work.
-
Pro plan’s daily video limit and continuity issues
By
–
Now here's what I'm not going to do: pretend this was flawless. The Pro plan currently caps video generations at 3 per day. When prompting a continuation clip, the model changed character outfits and altered the first character's appearance mid-conversation. Whether that's a
