the new Fable still can't tell a joke i think jokery evals plateaued with Google's PaLM models in 2022, no one has pushed SOTA since then maybe another 10 trillion parameters will do the trick!
RESEARCH
-

Tackling Anthropic’s 319-page System Card
By
–
So far, the minimum we need to know. Now it's time to tackle the System Card of only 319 pages 🙂 Link https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf …
-

Fable 5 / Mythos 5 most capable AI, twice Opus price, legal benchmark issue
By
–
It’s about twice as expensive as Opus, but Fable 5 / Mythos 5 is the most capable AI on the planet (Also, why does every model seem to suck on the legal agent benchmark?)
-
GPT-5.6 Cannot Compete with Fable 5 According to Evals
By
–
I mean, i guess so. obv. i havent tested both, or even fable sufficiently. but reading their blogpost and looking at the evals, i dont see how GPT-5.6 can compete right now. Of course, you wont use Fable 5 for every task and GPT-5.5 is a beast. but given the different use cases,
-

Improvement of vision models on Pokémon Red
By
–
The model also improves on vision tasks. Very interesting what they comment: where previous models could not surpass Pokémon Red in an agentic way, even by supplementing their vision with additional information and tools,
-

Claude 5 Fable: state-of-the-art on benchmarks, excels on longer and complex tasks
By
–



Claude 5 Fable tl;dr – It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research -The longer and more complex the task, the larger Fable 5’s lead over our other
-

Spectacular slope of the precision/cost frontier
By
–
By observing the precision/cost frontier on the FrontierCode benchmark we discussed, we can see the spectacular slope of computation at the time of testing this new model. Look how it rises! Even in the low configuration, the model performs — and costs — more.
-

Big leap in agentic programming and new FrontierCode Diamond benchmark
By
–
Regarding the first benchmark table they share, the big leap is observed especially in agentic programming (where labs get more value from users). The new FrontierCode Diamond benchmark released yesterday, designed to be more difficult than
-

Anthropic Releases Claude Fable 5, Mythos-Class Model After Hype
By
–
Breaking: @AnthropicAI just released Claude Fable 5: a Mythos-class model that has been the talk of the AI town for the past four months. Its capabilities have been teased relentlessly over this period. We can finally take it for a spin and see for ourselves if the hype matches
-

Mythos Launch: FrontierCode Benchmark, Opus 4.8 & GPT 5.5 Lack Scaling
By
–



Mythos is live! so excited to have our FrontierCode recognized as the next frontier coding bench. on FC Diamond, BOTH Opus 4.8 and GPT 5.5 don't meaningfully scale with effort, which many of you caught yesterday. Mythos/Fable posttraining have really applied that test time
