IME you still get a lot of control over this in the prompt itself. Even with the highest Pro settings, it thinks longer if you say to keep trying until some verification you propose passes etc.
@goodside
-
Critique of 4o’s Gen-Z personality and improvements in 5.x
By
–
Yeh that rings true to me. Also I always disliked the distinctly gen-z personality of 4o, and 5 felt like 4o’s personality invading o3. I notice this problem less with each successive 5.x release.
-
No malice assumed: TiKZ unicorns training not in OAI’s interest
By
–
I don’t see any reason to assume malice here—even if increased training on TiKZ unicorns is the true explanation, you’d expect it to happen naturally given the impact of Bubeck’s work. More generally it’s not in OAI’s interest to “bechmaxx” weird tests and they surely know this.
-
Reconstructing chess board position as subtask of image gen
By
–
The point is that reconstructing the chess board position from that prompt was a ludicrous proposition until very recently, and it can now do it as a subtask of other, more traditional image gen work like knowing what FF6 or Chipotle looks like.
-
Reflection on o3 and initial 5 router fiasco
By
–
I would love to hear more on why you think this some time. I loved o3 (and posted about lot about it!) but initial 5 was IMO a dud, especially the router fiasco. I’m happy to admit there’s been great progress since then though.
-
Claude’s appeal for 10x coders acknowledged by aging user
By
–
I actually like Claude a lot btw, despite rarely tweeting about it. I use it almost as much as ChatGPT for routine questions on my phone. But it clearly shines brightest for 10x coder types and I’m getting too old for that.
-
Shift to image gen despite anti-AI art people
By
–
Idk. Maybe I’m just getting lazy. There was a time I didn’t do image gen at all as a rule, because I hated dealing with anti-AI-art people. But images is where the tweet-verifiable frontier is now; the rest is all high-context coding.
-
Three categories of LLM tasks: economic, differentiating, and tweetable
By
–
I have little to say about GPT-5.x for the same reason I’ve long had little to say about Claude 4.x—which is that every txt2txt LLM task is at most two: 1) A task with real economic value
2) A task that differentiates frontier models
3) A task you would gladly read in a tweet -
OAI researchers ask about beginner recursive self-improvements
By
–
OAI researchers: “So, what are some good recursive self-improvements for someone just getting into recursively self-improving?”
-
AI will handle philosophy, making human decline irrelevant
By
–
All philosophy on this subject will soon be done by AI so it’s not a big deal if we get worse at it, even terminally so.