Seeing Isn't Knowing Do VLMs Know When Not to Answer Spatial Questions (and Why)?
GENERATIVE AI
-
ElevenLabs previews on-device Text-to-Speech at Warsaw Summit
By
–
At the ElevenLabs Summit in Warsaw, we previewed on-device Text to Speech – a new model architecture that delivers human-level quality on limited hardware without an internet connection. pic.twitter.com/iZuztsIR9N
— ElevenLabs (@ElevenLabs) 2 juin 2026At the ElevenLabs Summit in Warsaw, we previewed on-device Text to Speech – a new model architecture that delivers human-level quality on limited hardware without an internet connection.
-

Annotations enable AI collaboration by selecting context for Codex queries
By
–
Annotations will be a new way to collaborate with AI: users can select any context from the docs and ask Codex any questions.
-

GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization
By
–
GPU Forecasters Language Models as Selective Surrogates for Kernel Runtime Optimization
-

New edition of ‘Rise of the Robots’ on AI and jobless future
By
–
The new edition of "Rise of the Robots: Technology and the Threat of a Jobless Future" is now available! I have extensively updated the book to cover the latest advances in generative #AI and robotics and to examine the future economic and job market implications of the
-

Crafter: Multi-Agent Harness for Editable Scientific Figure Generation
By
–
Crafter A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
-
Martin Scorsese tests FLUX AI for storyboarding publicly
By
–
Legendary filmmaker Martin Scorsese signed on last year as an adviser to Black Forest Labs, the German AI startup behind FLUX image models.
— The Rundown AI (@TheRundownAI) 2 juin 2026
On Tuesday he went public, testing the tool on a single scene during preproduction.
His use is narrow: storyboarding only, complementing… pic.twitter.com/Nrc8OqD1bLLegendary filmmaker Martin Scorsese signed on last year as an adviser to Black Forest Labs, the German AI startup behind FLUX image models. On Tuesday he went public, testing the tool on a single scene during preproduction. His use is narrow: storyboarding only, complementing
-

Representation Forcing for Bottleneck-Free Unified Multimodal Models
By
–
Most Unified Multimodal Models still generate images through a frozen VAE, which means perception and generation are not fully learned in one model. This paper fixes this by making the decoder first predict
-
Drop video into Codex to recreate animations in MagicPath
By
–
You can drop a video into Codex, tell it "recreate these animations in MagicPath" and it'll just do it pic.twitter.com/OYGG0etrfO
— Pietro Schirano (@skirano) 2 juin 2026You can drop a video into Codex, tell it "recreate these animations in MagicPath" and it'll just do it

