5/ One result stood out. In one test the model correctly named the experiment and explained the Stroop task in detail. Then it failed the task anyway. Knowing the rule did not help it follow the rule.
AI
-
GPT-4o accuracy falls sharply with more words and color mismatch
By
–
4/ Everything fell apart. GPT-4o scored 91% on 5 words. At 10 words it dropped to 57%. At 40 words it hit 15%. When the colors and words were mismatched, accuracy fell to near zero.
-
AI models ace short lists but struggle with longer ones
By
–
3/ Researchers ran this test on the top AI models. GPT-5, Claude Opus 4.1, Gemini 2.5, and others. On short lists they aced it. 90% and up. Then the researchers made the lists longer.
-
The smartest AIs fail a simple color test
By
–
The smartest AI models on Earth just failed a color test that a 6-year-old can pass. The longer the test went, the worse they performed. One went from 91% to 15%. Scientists have discovered a flaw in the way AI
-
US AI policy mess: Trump’s Executive Order and Commerce order criticized
By
–
Where is the US with respect to AI policy? It’s a mess. Let’s start with Trump’s Executive Order from June 2, and then turn to last night’s Commerce order; neither gets things right. It’s fantastic that Trump has issued an executive order that encourages AI companies to do some
-
Turn Claude into 20+ marketing & business specialists bundle
By
–
Turn Claude into 20+ different specialists for marketing & business. Install real expertise, not just prompts. Get my Claude skills bundle
-
The two most absurd weeks in AI history
By
–
The last two weeks in AI have been the most absurd period in the history of tech. I can't even keep up. And that's literally what I do for a living. → Anthropic launched Claude Fable 5, its most powerful model ever made public. At the forefront of
-
Sparse models of that size: 40-60 GB, fast; dense models too slow
By
–
Sparse models of that size use about 40-60 gb and are fast enough for me (but dense models of that size I agree there bare too slow)
-
Anthropic CEO discovers Chinese founder giving away architecture surpassing Claude for free
By
–
Le CEO d’Anthropic qui ouvre X (Twitter)
— Jouhatsu | AI Influence Operator (@Jouhatsu_ai) 13 juin 2026
et découvre un fondateur chinois d'IA à 20 milliards de dollars distribuer gratuitement, en 40 minutes, l'architecture exacte qui surpasse Claude https://t.co/tZDSjCaIKh pic.twitter.com/IXs4QWENGPThe Anthropic CEO opens X (Twitter) and discovers a Chinese AI founder worth $20 billion distributing for free, in 40 minutes, the exact architecture that surpasses Claude.
-
Adaline reads traffic, writes evals, assembles stronger agent candidates
By
–
Went in expecting another trace dashboard. Instead Adaline reads the production traffic nobody on the team has time to read, clusters it into real behaviors, and writes hundreds of fresh evals against them every day.
— Chubby♨️ (@kimmonismus) 13 juin 2026
Then it assembles the stronger agent candidates and hands… https://t.co/giNgHgeUtKWent in expecting another trace dashboard. Instead Adaline reads the production traffic nobody on the team has time to read, clusters it into real behaviors, and writes hundreds of fresh evals against them every day. Then it assembles the stronger agent candidates and hands