A few months ago, the best LLM scored 5% on the USA Math Olympiad test. Models have been rapidly improving. Today, Google Gemini 2.5 scored 49%, which is better than 75% of the people who took the test (roughly the top 250 students in the USA).
@erikbryn
-

LLMs Surpass Benchmarks: Converting AI Capabilities Into Business Value
By
–
LLMs are blowing through benchmarks faster and faster. Next up, converting capabilities into business value.
-
Government cuts to life-saving research receive public support
By
–
I'm sorry to read this and all the other cuts. Supporting work like yours is one the government activities with the highest return. We should be doing *more*, not less. I'm incredibly grateful that you've devoted your life to this kind of life-saving research. Hang in there!
-
Anthropic Economic Index Launch with Alex Tamkin
By
–
I love what @AnthropicAI has done to create its Economic Index. No one understands it better than @AlexTamkin
. Join us on Monday and bring your questions! -
New Research Podcast Justified Posteriors Explores AI Papers
By
–
Check out this amazing new podcast by two of my amazing former postdocs: @AndreyFradkin and @SBenzell The format is that they state their priors, read a paper, discuss it on the pod, and then update their priors. It's called "Justified Posteriors" Link in next post…
-

AI Models Creating Original Research as Capabilities Advance
By
–
Interesting hypothesis that we may soon see an explosion of original work from AI models as they approach the capabilities of the best human researchers and writers.
-
AI Rapidly Improving Across Benchmarks: 11 Key Trends
By
–
AI is getting better at a bunch of benchmarks very, very fast; 11 more trends summarized here: https://
hai.stanford.edu/ai-index/2025-
ai-index-report
… -
David Autor discusses technology and jobs impact
By
–
I can't wait to see my friend and old MIT colleague @davidautor next Monday. His research about technology and jobs is always insightful. Please join us!
-
Bridging AI capabilities gap in enterprise business value delivery
By
–
Here's a terrific article by @Steve_Rosenbush in the @WSJ quoting me and others about how to address the gap between AI's amazing capabilities and the frustration many businesses have in delivering value.
-
Recognition for Innovation and Tax Policy Research Work
By
–
Congrats to @S_Stantcheva for being recognized for her terrific work on innovation and tax policy: