I also reviewed the three original responses for factuality, finding no clear errors from o1 or o1 pro, but one major one from o3-mini-high: (The film features no battles with Angels or Unit-01)
LLMS
-
o3-mini-high responses show Shinji, Asuka mention counts
By
–
n=1 per model above but to verify for this post I generated 5 more responses from o3-mini-high, yielding from each only 4-7 mentions of Shinji, 0-4 mentions of Asuka, and 0 total mentions of Rei, Misato, and Gendo.
-

Word counts and character mentions in AI responses to End of Evangelion
By
–
Word count and character name mentions for responses (n=1 per model) to the prompt "Explain End of Evangelion" for ChatGPT o3-mini-high, o1, and o1 pro
-
Plot testing o3-mini’s world knowledge degradation compared to o1
By
–
I made this plot as a quick vibe test for how "mini" o3-mini's world knowledge is. You can make small heuristic evals for domains you care about to see how much o3-mini's degradation vs. o1 matters to you — also, to measure your progress prompting back to parity.
-
Lex Fridman Praises o3-mini-high, Anticipates o3 and o3pro Release
By
–
Congrats, I'm enjoying o3-mini-high a lot, well done Can't wait for o3/o3pro!
-

Impressive Results Switching AI Agents from o1 to o3-mini
By
–
Incredible first results after switching AI agents from o1 to o3-mini!
-

OpenAI o3-mini Surpasses o1-mini in Codeforces Competitive Programming
By
–
On Codeforces competitive programming, @OpenAI o3-mini achieves progressively higher Elo scores with increased reasoning effort, all outperforming o1-mini. With medium reasoning effort, it matches o1’s performance.
-
AI Unlocking Ancient Texts and Rewriting History
By
–
How AI is unlocking ancient texts — and could rewrite history
#AI #AIio #AIInnovation #ML #DataScience #Futureofwork @Scobleizer @AndrewYNg @drfeifei @KirkDBorne @fchollet @rowancheung @antgrasso -
Optimizing AI Reasoning with Shorthand Chain of Thought Language
By
–
I asked o3 on how I’d optimize this for faster reasoning and one of the ideas it gave me was to invent a shorthand language for the chain of thought to guzzle less tokens https://
gist.github.com/willccbb/46767
55236bb08cab5f4e54a0475d6fb
…

