8/ It fits a pattern researchers keep hitting. The models that pass medical exams and win math olympiads still fall apart on simple tasks once those tasks run long. Source: Patel, Wang, Fan. PNAS Nexus, 2026.
MACHINE LEARNING
-
Color game reveals models’ inability to maintain focus over long input
By
–
7/ This is why a color game matters. The models still had the knowledge. What they lacked was the ability to hold focus and resist a pull across a long stretch of input. That gap is the finding.
-
AI models default to reading over naming colors
By
–
6/ The cause is in how these models are built. They were trained on text, so reading words is their strongest instinct. Naming the color means fighting that instinct. A human brain can suppress the urge. The model defaults to reading.
-
GPT-4o accuracy falls sharply with more words and color mismatch
By
–
4/ Everything fell apart. GPT-4o scored 91% on 5 words. At 10 words it dropped to 57%. At 40 words it hit 15%. When the colors and words were mismatched, accuracy fell to near zero.
-
AI models ace short lists but struggle with longer ones
By
–
3/ Researchers ran this test on the top AI models. GPT-5, Claude Opus 4.1, Gemini 2.5, and others. On short lists they aced it. 90% and up. Then the researchers made the lists longer.
-
The smartest AIs fail a simple color test
By
–
The smartest AI models on Earth just failed a color test that a 6-year-old can pass. The longer the test went, the worse they performed. One went from 91% to 15%. Scientists have discovered a flaw in the way AI
-
The two most absurd weeks in AI history
By
–
The last two weeks in AI have been the most absurd period in the history of tech. I can't even keep up. And that's literally what I do for a living. → Anthropic launched Claude Fable 5, its most powerful model ever made public. At the forefront of
-
Sparse models of that size: 40-60 GB, fast; dense models too slow
By
–
Sparse models of that size use about 40-60 gb and are fast enough for me (but dense models of that size I agree there bare too slow)
-
Anthropic CEO discovers Chinese founder giving away architecture surpassing Claude for free
By
–
Le CEO d’Anthropic qui ouvre X (Twitter)
— Jouhatsu | AI Influence Operator (@Jouhatsu_ai) 13 juin 2026
et découvre un fondateur chinois d'IA à 20 milliards de dollars distribuer gratuitement, en 40 minutes, l'architecture exacte qui surpasse Claude https://t.co/tZDSjCaIKh pic.twitter.com/IXs4QWENGPThe Anthropic CEO opens X (Twitter) and discovers a Chinese AI founder worth $20 billion distributing for free, in 40 minutes, the exact architecture that surpasses Claude.
-
Adaline reads traffic, writes evals, assembles stronger agent candidates
By
–
Went in expecting another trace dashboard. Instead Adaline reads the production traffic nobody on the team has time to read, clusters it into real behaviors, and writes hundreds of fresh evals against them every day.
— Chubby♨️ (@kimmonismus) 13 juin 2026
Then it assembles the stronger agent candidates and hands… https://t.co/giNgHgeUtKWent in expecting another trace dashboard. Instead Adaline reads the production traffic nobody on the team has time to read, clusters it into real behaviors, and writes hundreds of fresh evals against them every day. Then it assembles the stronger agent candidates and hands