Claude-Fable-5 on BullshitBench: it refused a THIRD of questions (with multiple attempts). Excluding refusals, it performed well, worse than some other Claude models (chart below).
LLMS
-
Shower thoughts: autoencoders as good addition after LoRA
By
–
This was purely based on shower-thoughts :D.
As someone called out, autoencoders would have been a good addition after LoRA. -
Fable 5: Leap Forward, 100% Coding in Terminal
By
–
Fable 5 represents the greatest leap forward I have felt in our models since Opus 4.5 last November. After the release of 4.5, I uninstalled my IDE when I realized that I had done 100% of my coding in a terminal for a few weeks. With Fable,
-
Fable praised for intelligence and distinctive replies
By
–
Fable is really smart and replies very differently from other models
-

Anthropic’s guardrail concerns versus model unusability regrets
By
–
I understand that Anthropic's concerns about the model being misused without guardrails are significant. And I take that seriously. We're talking about a technology with unforeseen potential. However, the fact that it was, in some cases, literally unusable is regrettable.
-

A model more capable than a year ago, priced below Opus 4.1
By
–
If you really look at it, this model, much more capable than what we had a year ago, is released with a price below Opus 4.1.
-
The importance of self-verification loops in the era of powerful models
By
–
We talk a lot about how important it is to set up self-verification loops. Especially in the age of powerful models that can run for long periods of time, self-verification is a key ingredient that enables the model to run for much longer, delivering a result that is closer to… https://t.co/NHiral0F9j
— Boris Cherny (@bcherny) 9 juin 2026We talk a lot about the importance of setting up self-verification loops. Especially in the era of powerful models that can operate for long periods, self-verification is a key ingredient that allows the model to function much more
-
LLM shadow-banning was not planned for 2026
By
–
The shadow-banning of LLM development was certainly not on my bucket list for 2026.
-
Limiting Claude for cutting-edge LLM development widens the gap
By
–
“limiting Claude’s effectiveness for requests aimed at developing cutting-edge LLMs” And that’s how they plan to widen the gap even further
-

34 days from signing deal to Mythos-class model GA launch
By
–
for those keeping track at home it was 34 days between signing this deal and launching Mythos-class model GA to the world. https://
x.com/leerob/status/
2052059466821198061?s=20
… building on @nvidia stack means you can just do things™.
