Codex: "Study all of the conversations we've ever had (not just scan the memories), then use the imagegen skill to generate a perfect joke about me; make the image clean and sharp and in a unique style"
MULTIMODAL AI
-
Image V2 AI Transforms Graphics Generation and Recontextualization
By
–
Now, what a wild thing this image v2 is for updating and recontextualizing images and graphics from others on the fly
-
Grok Voice Think Fast 1.0 State-of-the-Art Voice Model Launch
By
–
Introducing Grok Voice Think Fast 1.0 A state-of-the-art voice model built for complex, multi-step workflows with snappy responses and high accuracy. It takes the top spot on the Tau Voice Bench and handles real-world messiness like noise, accents, and interruptions better than
-

Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
By
–
Fine-Tuning Diffusion Models via Intermediate Distribution Shaping paper: https://
huggingface.co/papers/2510.02
692
… -

GPT-5.5 Available via Codex API Backdoor for Image Generation
By
–

GPT-5.5 may not be in the official OpenAI API… but it's available via the apparently approved-of Codex API backdoor So I used that to make these pelicans (default and xhigh)! https://
simonwillison.net/2026/Apr/23/gp
t-5-5/
… -
Early Access to GPT-5.5 Pro Version Announced
By
–
I had early access to GPT-5.5. It is very good, especially the Pro version. Full writeup very shortly.
-

GPT-5.5 Represents New Class of Artificial Intelligence
By
–
GPT-5.5 is a new class of intelligence.
— Greg Brockman (@gdb) 23 avril 2026
This intelligence makes it intuitive to use; it completes challenging tasks with little micromanagement. Also very token efficient, and runs with low latency and at scale.
A real step toward a new way of getting computer work done. https://t.co/xOHtb72h22GPT-5.5 is a new class of intelligence. This intelligence makes it intuitive to use; it completes challenging tasks with little micromanagement. Also very token efficient, and runs with low latency and at scale. A real step toward a new way of getting computer work done.
-

Vision Banana: Rethinking AI Model Generalization in Vision
By
–
Vision Banana: Rethinking How AI Models See and Generalize
— Satya Mallick (@LearnOpenCV) 23 avril 2026
In this episode of Artificial Intelligence: Papers and Concepts, we explore Vision Banana, a concept that challenges how vision models learn and generalize from visual data. Instead of focusing purely on performance… pic.twitter.com/vgp0LQbMFiVision Banana: Rethinking How AI Models See and Generalize In this episode of Artificial Intelligence: Papers and Concepts, we explore Vision Banana, a concept that challenges how vision models learn and generalize from visual data. Instead of focusing purely on performance
-
Gemini 3.1 TTS Introduces Audio Tags for Vocal Style Control
By
–
Last week, we launched Gemini 3.1 TTS, our latest and best text-to-speech model. This new model introduces [awe] audio tags, an intuitive way to guide vocal style, pace, and delivery.
— Google AI (@GoogleAI) 23 avril 2026
Here are some tips on the best ways to use audio tags in your prompts:
1. All inline tags must… pic.twitter.com/YDbBLs5DcpLast week, we launched Gemini 3.1 TTS, our latest and best text-to-speech model. This new model introduces [awe] audio tags, an intuitive way to guide vocal style, pace, and delivery. Here are some tips on the best ways to use audio tags in your prompts: 1. All inline tags must
-
ChatGPT Pro vs. Thinking: Image quality comparison and model distinctions
By
–
Ran with latest ChatGPT Pro with Extended mode. I believe the final model (ChatGPT Images 2.0) is distinct. It seems clear to me you get better images with Pro than Thinking, but the details of how they work together isn’t fully explained anywhere AFAIK.
