minWM
— AK (@_akhaliq) 29 mai 2026
A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models pic.twitter.com/hxUocDWGJj
minWM A Full-Stack Open-Source Framework for Real-Time Interactive World Models
By
–
minWM
— AK (@_akhaliq) 29 mai 2026
A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models pic.twitter.com/hxUocDWGJj
minWM A Full-Stack Open-Source Framework for Real-Time Interactive World Models

By
–
Multi-model person-of-interest ID + weapon detection, built on the Voyager SDK. The "weapon" was a lightsaber (Count Dooku's hilt). The specs were real though! 3x 4-chip Metis cards, 48 AIPU cores → 2.5 PetaOPS, 1,440+ model inferences/sec across multiple 8K streams at
By
–
Larus went ham with this one! Love the synced highlighting on the camera path, something I wanted to try myself.
— Bilawal Sidhu (@bilawalsidhu) 29 mai 2026
Makes me think these could end up as spatial reasoning benchmarks for ai video models, esp in cities with existing 3d data as ground truth. pic.twitter.com/x8BBPPuOEu
Larus went ham with this one! Love the synced highlighting on the camera path, something I wanted to try myself. Makes me think these could end up as spatial reasoning benchmarks for ai video models, esp in cities with existing 3d data as ground truth.
By
–
Introducing Wayve Labs: @wayve_ai
's frontier research lab for physical AI. Wayve Labs is where we'll pursue pioneering research in world models, spatial intelligence, and much, much more. Thanks to @ryajetha for the great write-up
By
–
Go behind the scenes to learn more about how The Rogue was made in under a month, by a single person with Runway.
— Runway (@runwayml) 29 mai 2026
The Rogue is part of Project Luxo: a new initiative exploring how AI-generated video has crossed the uncanny valley. pic.twitter.com/HlB4Ec8Cd0
Go behind the scenes to learn more about how The Rogue was made in under a month, by a single person with Runway. The Rogue is part of Project Luxo: a new initiative exploring how AI-generated video has crossed the uncanny valley.
By
–
Object detection is shifting from "models that recognize fixed categories" to "models that understand concepts described in language."
— Satya Mallick (@LearnOpenCV) 29 mai 2026
YOLOE delivers open-vocabulary detection at full YOLO speed — text module fused into the head, zero runtime overhead.
Full tutorial + code:… pic.twitter.com/0yRPYv5jRU
Object detection is shifting from "models that recognize fixed categories" to "models that understand concepts described in language."
YOLOE delivers open-vocabulary detection at full YOLO speed — text module fused into the head, zero runtime overhead.
Full tutorial + code:
By
–
X Square Robot presents WALL-WM — a World Action Model that uses semantic events as atomic units, aligning language, video, and action granularities. https://t.co/JAKBdwnCKE
— 机器之心 JIQIZHIXIN (@jiqizhixin) 29 mai 2026
X Square Robot presents WALL-WM — a World Action Model that uses semantic events as atomic units, aligning language, video, and action granularities.

By
–
Google’s new anything-to-anything #AI model is wild
by Allison Johnson @verge Learn more: https://
bit.ly/49iHZqj #GenerativeAI #ArtificialIntelligence #MachineLearning #ML

By
–
How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test of combining research, code, design and stats for an AI. https://
veil-of-history.netlify.app