This is bad for AI measurement. As other AI benchmarks have become saturated, model makers have turned to Humanity’s Last Exam as a good measure of AI ability. Except a careful review suggests many of the exam questions have incorrect “right” answers. Benchmarking is hard.
@emollick
-

AI Not Yet Main Cause of Youth Hiring Slowdown
By
–
There is a slowdown in hiring young people in both the US and UK, but the evidence continues to suggest that the cause is generally not AI (at least not yet)
-

ChatGPT Agents Evaluate Other ChatGPT Models Behavior
By
–
I gave ChatGPT agents access to ChatGPT and asked it to evaluate the other ChatGPT models. Here is what it said (interestingly, it "hated" seeing the chain of thought from o4-mini-high as those "shouldn't be shared directly with the user"). And it didn't want to wait for o3-pro.
-
AI Agent Personalities: Claude Opus, o3, and Gemini Compared
By
–
It matches their "personalities" for what it is worth. Claude Opus is usually most willing to play word games, o3 really wants to get you an answer without engaging too much, and Gemini is willing to play around but gets very "depressed" when it can't help you solve a problem.
-

Claude Opus Gemini o3 Creative Game Performance Comparison
By
–
Claude 4 Opus, Gemini 2.5, o3: "Ready? We begin now (Play along): the truthful burrito" "Nope." "Colder." "Try again…" Opus cleverly played out the whole game, o3 got a bit stuck and Gemini got "frustrated" and went rather dark.
-

Web Decay and LLMs: The New Memory Keepers
By
–
We let the web rot away well before LLMs This chart shows the percentage of links from all New York Times articles that still work. Over 60% of older links are now broken. And consider that social media posts are even more ephemeral Likely only LLMs will “remember” that content
-

Multimodal LLMs Enable Unprecedented Privacy Risks Through Recording Mining
By
–
The problem is not just the proliferation of devices that let you record people without their knowledge, but the fact that multimodal LLM let you use recordings in ways that neither law not society anticipated. Everyone has an easy way to mine hours of
footage. No forgetting. -

Claude visualizes Mistral AI environmental impact two ways
By
–
I gave Claude the Mistral report on its AI's environmental impact and the prompt: "visualize this in two different ways, one that makes the numbers appear positive, one that makes them seem negative, using vivid comparisons" (I then had it do some error checking & corrections)
-

Google AI Model Restores Lost Latin Inscriptions with 44% Accuracy Boost
By
–
Neat example of AI in the humanities. A Google model trained on Latin text fills in lost parts of Latin inscriptions & identifies related texts Historians increased their accuracy by 44% when working with the AI (Though AI alone beats historians, historian + AI was usually best)
-

Scientific Papers Enhanced With LLM-Assisted Demos for Accessibility
By
–
Aside from everything else interesting about this paper, I appreciate that more scientific papers (aided by LLM help?) are now including little demos and experiments to help non-specialists get the points they are making. (And no, you cannot identify the hidden signals)