Even the image models are not good enough world models yet, gpt-image-2 and nano banana make pretty obvious world model like mistakes. is way harder, so no hope of that anytime soon. Maybe with 2-3 OOMs it gets good enough, but that doesn't feel like a reasonable thing to
Even image models fail as world models; video is harder
By
–