I came up with a somewhat foolish new benchmark for testing image generation models, to exercise the new ChatGPT Images 2.0: "Do a where's Waldo style image but it's where is the raccoon holding a ham radio"
PROMPT ENGINEERING
-
AI complexity trade-offs: coding versus daily quick tasks
By
–
actually makes sense, especially for more complex tasks like coding vs for daily quick usage (iterating on some basic research)
-
Speed-typing with AI: Casual Communication Patterns
By
–
does anyone else "speed-type" when talking to AI? ignor spellign mistakes skip commas use agro abbrev etc
-
AI Storytelling: Building Real-World Information Scripts
By
–
Oh, I get the worlds. But I want to be able to have a corner of one of them to be a place people can learn what's happening in the real world. 🙂 Like a new kind of speech where an AI is telling people about the world. Or to be able to have an AI, or me, build a script
-
Forcing Extended Thinking in Claude Opus 4.7 Image Generation
By
–
This piece thinks that the reason I got a crap pelican riding a bicycle from Opus 4.7 is that it didn't think about it first, I'm trying to figure out if there's a way to force it to think that I've missed
-
Claude Opus 4.7 Extended Thinking API Budget Tokens Availability
By
–
Claude Opus 4.7 with adaptive thinking via the API… am I missing something or is it not possible any more to force it to think? (Prompt hacks like "think step by step" don't count here, I mean the equivalent of budget_tokens or effort: high in previous Claude models)
-

OneVL: One-Step Latent Reasoning and Planning with Vision-Language
By
–
OneVL One-Step Latent Reasoning and Planning with Vision-Language Explanation paper: https://
huggingface.co/papers/2604.18
486
… -
Opus 4.7 Max Shows Regression in Problem Solving Tasks
By
–
Slight regression with Opus 4.7 (Max) though, my guess is that the 'improved instruction following' + lots of thinking about how to solve the problem isn't helping Opus in this case