OpenAI Deep Research achieves 26.6% on Humanity’s Last Exam — more than double prior best o3-mini-high at 13.0% Note Deep Research is browsing + Python vs. pure LLMs for others so not totally comparable, but by design Q’s are Google-poof to non-experts so still very impressive
OpenAI Deep Research Aces Humanity’s Last Exam, Outperforming Previous Best
By
–
