What has been expected so far is a possibility for 4o to generate images with multimodality and not by talking to DALL-E model as it user to be before. Cuz earlier, 4o was generating a prompt and was passing it to DALL-E
Technical discussion on GPT-4o native multimodal image generation
By
–