GPT IMAGE 2 is wild.. Try asking it to create an image based on our conversation history so far.
MULTIMODAL AI
-

Where’s Waldo in 3D: New Image Generator’s Impressive Detail Capabilities
By
–
Where's Waldo in 3D: crazy detail from the new image generator Source: https://
reddit.com/r/ChatGPT/comm
ents/1st5qnx/wheres_wally_3d_crazy_detail_new_img_gen/
… -
Images 2.0 Achieves Major Qualitative Threshold Breakthrough
By
–
Images 2.0 really got over some important qualitative threshold for me that I didn't know existed.
-
ChatGPT Images 2.0: Studio-Quality Designs From Simple Prompts
By
–
ChatGPT Images 2.0 is getting out of hand. Designers are getting worried. People are generating studio-quality designs from a single prompt. 10 wild examples:
-
Vision Banana: Image Generation Improves Generalist Vision Learning
By
–
Wow! Vision banana confirms that image generators are great generalist vision learners. Pretraining alone gets you zero shot segmentation, depth, normals etc – while beating out specialist models. If you teach a model to draw – you also teach it to see! pic.twitter.com/FGsyzxlrmW
— Bilawal Sidhu (@bilawalsidhu) 23 avril 2026Wow! Vision banana confirms that image generators are great generalist vision learners. Pretraining alone gets you zero shot segmentation, depth, normals etc – while beating out specialist models. If you teach a model to draw – you also teach it to see!
-

Image Generators Emerge as Generalist Vision Learning Models
By
–
Image Generators are Generalist Vision Learners Paper: https://
arxiv.org/abs/2604.20329
Project: https://
vision-banana.github.io -

Google Vision Banana: Instruction-Tuned Generalist Image Generator Model
By
–
Huge! Google just proved Image Generators are Generalist Vision Learners! They introduce Vision Banana, a model built by instruction-tuning a base image generator (Nano Banana Pro). Instead of using special systems for different tasks, they reframe every vision problem—like
-

ChatGPT Images 2.0 Review: Best Model Currently Available
By
–
NEW VIDEO in the LAB! Here I bring you a video testing, experimenting with, and judging ChatGPT Images 2.0's new image model The best model currently?
Well, I'll save you from watching the video: Yes, it is. Although it has a problem… Link below -
Advancements in Generative World Models and LTX-2 Optimization
By
–
“Every pixel will be generated”
— Linus ✦ Ekenstam (@LinusEkenstam) 23 avril 2026
— Jensen Huang.
World models have been streaming for some time but they have a temporal cut-off at very short session times. 60-120 seconds. Then done.
This approach is different, and heavily optimized version of LTX-2 has been used to make it… https://t.co/3Xdm2c6lUh“Every pixel will be generated” — Jensen Huang. World models have been streaming for some time but they have a temporal cut-off at very short session times. 60-120 seconds. Then done. This approach is different, and heavily optimized version of LTX-2 has been used to make it
-

LLaDA2.0-Uni: Unified Multimodal Understanding and Generation with Diffusion
By
–
LLaDA2.0-Uni Unifying Multimodal Understanding and Generation with Diffusion Large Language Model paper: https://
huggingface.co/papers/2604.20
796
…
