Current agents only do 30% of complex real company tasks in this paper. Though note benchmarks are a floor, not a ceiling, if:
1) More recent models show improvement in the benchmark, suggesting future models may do it
2) Better prompting/tools would make the AI perform better.
@emollick
-

Current AI Agents Complete Only 30% of Complex Tasks
By
–
-

Disagreement Dial: Calibrating AI Model Confidence Levels
By
–
A disagreement dial would be a really useful thing for AI models. A slider that goes from “I am always right” to “I am always wrong” with intermediate steps.
-
Democratizing Education Through Technology: Opportunities and Social Media Risks
By
–
As someone who has been working on democratizing education for decades (Media Lab, Wikipedia admin, making early MOOCs, games for education, AI and teaching), I do see every advance making a difference, so I am still very positive But social media remains a warning worth heeding
-
Internet’s Educational Promise: Why Access Doesn’t Equal Impact
By
–
The fact that people use the internet mostly for entertainment isn't a weird or surprising But you also have access to courses on every topic by experts, every major out-of-copyright book, can talk to people from anywhere, etc. The impact of that is smaller than I once expected.
-
Web’s Promise: Universal Access Failed to Bridge Divides
By
–
It really is not what most people who was working on building the early web in the late 1990s were expecting. Universal access to information was going to transform everything, creating widespread learning and bridging divides. It really is shocking how much that didn't happen.
-
Why AI and Information Access Don’t Fix Misinformation
By
–
X (and other social media sites) make our 1990s optimism about the Information Age seem silly. Even with all of the world's information a click away (& a free AI that can help explain that information in a personalized way), half-mangled anecdotes with no source win every time.
-
General Purpose Technologies: Unpredictable Impacts on Society
By
–
I think it should be obvious, but worth repeating that General Purpose Technologies have many effects, good and bad, that are not predictable to the technology's creators and which impact culture, society, and the economy in complex ways that are impossible to fully anticipate.
-

AI Models Replace Google for Complex Multi-Constraint Searches
By
–
An example of the type of search (would require reading multiple sites, balancing multiple constraints) where o3/Gemini 2.5 Pro has completely replaced Google for me.
-
Enterprise Adoption of Agentic Systems Slower Than Expected
By
–
I would be very surprised, given what I have seen in companies, if "agentic systems" are actually agentic in the way that many on X interpret that term. Firms are still adopting chatbots slowly, I do not think most of them have moved on to agents in any real way.
-
AI Education Misinformation: Debunking Common Misinterpretations
By
–
This interpretation of the paper has joined the pantheon of AI misinformation that I get asked about often. There are legitimate worries about AI use in education, just not this one.