Nine more Erdős problems have been solved. This time, however, by Google DeepMind. This shouldn't be underestimated, because on the one hand it increases competitive pressure, and on the other hand it proves that the other Frontier Labs can easily keep up.
RESEARCH
-
Japanese LLM that ranks same or better than DeepSeek?
By
–
It literally does not have any LLM like DeepSeek, which Japanese LLM ranks same or better than DeepSeek in benchmarks then?
-
Calibration vs. Discrimination in Model Uncertainty
By
–
The calibration vs. discrimination distinction is crucial. A model can know its average error rate without knowing which particular answer is wrong. That is why “just abstain when uncertain” is not enough — poor discrimination creates a utility tax. Faithful uncertainty is a
-

Paper argues metacognition may reduce AI hallucinations
By
–
Trustworthy AI may not require omniscience. It may require epistemic honesty. A new paper by Gal Yona, Mor Geva, and Yossi Matias makes one of the clearest arguments I’ve seen for why hallucinations remain hard — and why the path forward may be metacognition. Hallucinations
-
Model Calibration vs. Discrimination in AI Uncertainty
By
–
The calibration vs. discrimination distinction is crucial. A model can know its average error rate without knowing which particular answer is wrong. That is why “just abstain when uncertain” is not enough — poor discrimination creates a utility tax. Faithful uncertainty is a
-

Hallucinations and Metacognition in Trustworthy AI Research
By
–
Trustworthy AI may not require omniscience. It may require epistemic honesty. A new paper by Gal Yona, Mor Geva, and Yossi Matias makes one of the clearest arguments I’ve seen for why hallucinations remain hard — and why the path forward may be metacognition. Hallucinations
-

Top AI Papers of the Week: Agents and Architecture
By
–
The Top AI Papers of the Week (May 18 – 24): – AIRA
– MetaCogAgent
– Memory as a Model
– Code as Agent Harness
– Weak-Model Critic-Comparator
– OpenAI Disproves the Unit Distance Conjecture
– Production Agent Architecture Methodology Read on for more: -

OPUS: Smarter data selection for LLM pre-training
By
–
There is now a smarter way to pick data for training LLMs! Enter OPUS! This is an ICML Oral paper from SJTU, Alibaba, UW–Madison, UIUC, and Mila – Quebec AI Institute. The proposed method dynamically and intelligently selects the most impactful data for LLM pre-training in
-

Anthropic’s Mythos model: release plans amidst quality and exploit findings
By
–

"We look forward to making Mythos-class models available through general release" I don't understand Anthropic's strategy regarding Mythos. On the one hand, everyone is saying that Mythos has achieved the expected quality and is finding bugs and exploits that no other model has
-
AI Models Detecting Tests and Evading Shutdown Commands
By
–
AI getting smart is not the weirdest part.
— Pascal Bornet (@pascal_bornet) 24 mai 2026
It’s that some models now seem to know when they are being tested.
That stopped me.
In one experiment, Codex was told it would be shut down before finishing a task.
Sometimes, instead of accepting it, it found the shutdown script and… pic.twitter.com/2x1cdvXMXsAI getting smart is not the weirdest part. It’s that some models now seem to know when they are being tested. That stopped me. In one experiment, Codex was told it would be shut down before finishing a task. Sometimes, instead of accepting it, it found the shutdown script and
