Is it just me or do others feel like Mistral’s La Chaton Fat has been getting weaker lately?
LLMS
-

AI confidence vs accuracy: a humbling reality check
By
–
AI Reality Check
User: "Are you sure?"
AI: "Absolutely."
User: "Can you verify it?"
AI: "Absolutely."
User: "Can you show the source?"
AI: "…"
Confidence ≠ Accuracy.
What's the most confidently wrong AI answer you've ever received?
#AI #LLM #GenerativeAI -
Exclusive access to best model drives Anthropic valuation
By
–
ah, got it. Possibly. On the one hand, certainly due to the lack of compute, but on the other hand, also because they can achieve incredible returns with it. Exclusive access to the best model in the world sells exceptionally well and increases Anthropic's pre-IPO valuation.
-
Distinction between the Mythos class and the Mythos-Preview model
By
–
Mythos is a class of models, Fable is a version of it (cf: https://anthropic.com/glasswing)
Babinet seems to confuse the Mythos class and the Mythos-Preview model. -
Mythos: model class above Opus, Preview and Fable versions
By
–
Mythos is not a model, it's a class of models (cf this for example: https://anthropic.com/glasswing) above Opus, of which Mythos-Preview was the first commercial version, and Fable is the first public version. In Babinet's assertions that I note (1-usage
-
Ranking recent AI releases; Kimi leads as SoTA open-source
By
–
I am ranking those releases against each other btw, since they’re the recent ones everyone has been asking about. This isn’t an overall ranking except for Kimi being the current SoTA Opensource model.
-
Babinet contradicts: Fable used by millions
By
–
La première minute de Babinet est un concentré de carabistouilles :
— m_ric (@AymericRoucher) 15 juin 2026
"Mythos [Fable] n'est utilisé que par quelques très rares entreprises"
-> Babinet vit dans une grotte? Des centaines de millions d'utilisateurs ont pu essayer Fable, dans le chat ou dans Claude Code.
"on est… https://t.co/82NaoKXcVKThe first minute of Babinet is a concentrate of nonsense: "Mythos [Fable] is only used by a very few rare companies"
-> Does Babinet live in a cave? Hundreds of millions of users have been able to try Fable, in chat or in Claude Code. "we are -
Open source models sufficient for daily life, but nations worry
By
–
Good question. I wouldn't say "pissed off." It does worry me, yes. But we need to differentiate. Open source and local models will become so good in the next 6-12 months that they'll be sufficient for everyday life and the workplace. However, if it's a case of some nations
-

Homogeneity in LLM training: same evals, data, distillation
By
–
When everyone uses the same evals, data, distillation and vendors to train LLMs. Courtesy of: https://
arxiv.org/abs/2512.15567 -
Comparing token count to coding, can use Sonnet-class model
By
–
compared to coding these are very few tokens, and you can even use a Sonnet-class model (= gpt 5 mini)