The artificial analysis index is a normalized score of several benchmarks (and has changed over time) it is fine for roughly comparing models, it is not useful for trend analysis and it is unclear what individual point differences in the scores mean.
@emollick
-

Open vs Closed AI Models: Beyond Benchmark Gaps
By
–
This is a good explanation of why the gap between open and closed models is larger than it appears in benchmarks. I would add in that current open models are also more fragile than closed: they handle out-of-distribution problems far less well & have lower emergent capabilities.
-
Frontier Agent Benchmarking Struggles to Capture Real Progress
By
–
Its getting hard to benchmark frontier agent performance on longer tasks. Repeated measurement is very expensive and there are differences between using models in harnesses versus via APIs. I suspect benchmarks understate progress, they are built for models, not harnessed agents
-

AI Agents Drive Shift From Bubble Talk to Data Center Demand
By
–

I was quoted a couple times in this Atlantic article, but that isn’t (the only) reason I think it is good. It lays out the reasons why we whipsawed from “AI is a bubble” to “there are not enough data centers” in less than six months. Spoiler: its agents. https://
theatlantic.com/economy/2026/0
5/ai-bubble-revenue-anthropic/687022/
… -
AI Benefits Require Integration With Organizations Not Just Individuals
By
–
Organizations are already superhuman intelligences. The University of Pennsylvania or Walmart or whatever is far more capable than any human. That is why the focus on AIs as individual productivity tools hits a natural limit, many benefits of AI depend on integration with firms.
-

o1 Outperforms Doctors on Medical Benchmarks and ER Cases
By
–


New paper (on an old AI) tests o1 against doctors on medical benchmarks & real ER cases: “across a variety of scenarios and applications, the large language model outperformed both human physicians and older models” The potential suggests an “urgent need for prospective trials.”
-
GPT Image Gen Creates Progressive Grid Sequences From Dogs to Gatsby
By
–
GPT-imagegen-2: "make 5×5 grid of dog photos, where each photo gets noticeably cuter"
…now cats
…now man-eating squid
…now covers of the book the Great Gatsby -

AI Therapy Chatbot Improves Mental Health Outcomes in Mexican Women Trial
By
–



Randomized trial of an AI therapy chatbot on Mexican women found “improved mental health by 0.3 SD over 6 months with no evidence of an increase of severe cases; improved sleep, healthful behaviors, daily functioning & labor market outcomes” Big results for a cheap intervention.
-

Frequent AI Users Can Spot AI-Generated Prose by Its Tells
By
–
"Load bearing," "I keep coming back to," "Not X, but Y" A curse of using AI a lot is that you realize how much of the writing around you is just AI, now People who don't use AI have been unable to identify AI prose on sight, but those who use it a lot can spot the tells easily
-
Regulating Open-Source AI Models Poses Major Policy Challenges
By
–
For better or worse, regulation for closed-source models served by a few (quite large) companies is easy. It is not as easy to imagine how you regulate open-source models that can be served by a range of decentralized players. Suspect that will become a big policy discussion soon
