Didn't have enough tokens so it cut off at 'End World'
LLMS
-

Anthropic’s new Fable 5 safeguards quietly limit effectiveness
By
–
Anthropic’s new Fable 5 safeguards are fascinating. When the model is used for frontier LLM development, it apparently does not simply refuse or warn the user. Instead, it quietly limits its own effectiveness through techniques like prompt modification, steering vectors, and
-

Apple’s Core AI runs models entirely on-device
By
–
Apple finally did it. Its new framework, Core AI, runs models entirely on Apple silicon, so inference happens on the user's device with zero server calls and zero token bills. That means Qwen, Mistral, and SAM3 running natively across iPhone, iPad, Mac, and Vision Pro. It's a
-

Anthropic would limit capabilities to maintain competitive advantage
By
–
Pretty crazy this that is being shared where Anthropic would be limiting the model's capabilities when used to improve and create better LLMs. They sell it as a security measure but it is clear that they do it to maintain their competitive advantage.
-

OSCAR: 2-bit KV cache for LLMs without accuracy loss
By
–
Can LLMs run on ultra-low-bit memory without tanking accuracy? Researchers from Together AI, University of Sydney, and UIUC present OSCAR — a method that uses offline, attention-aware covariance analysis to design fixed rotations and clipping thresholds for 2-bit KV cache
-
My Hermes Agent using Claude Fable 5 to write git commit
By
–
My Hermes Agent using Claude Fable 5 to write git commit -m "fixed" pic.twitter.com/Xq5Vk7Oyvq
— Shubham Saboo (@Saboo_Shubham_) 9 juin 2026My Hermes Agent using Claude Fable 5 to write git commit -m "fixed"
-

Cost frontiers: more expensive and superior model, Fable surpasses Opus 4.8
By
–
The cost frontiers, as expected, show an upward and rightward shift, meaning we have a more expensive and notably superior model (the low of Fable remains well above the x-high of Opus 4.8)
-

Benchmark tables: the model dominates in everything evaluated
By
–

Here are the complete benchmark tables. The model basically dominates in everything evaluated in these tables (Humanity Last Exam, CriPT, ArxivMath, HealthBench, etc.)
-

Build Reliable GenAI Applications with AI Evals, Observability, and Testing
By
–
Workshop hosted by @PacktPublishing @PacktDataML — "Build Reliable GenAI Applications with AI Evals, Observability, and Testing" 𝗥𝗲𝗴𝗶𝘀𝘁𝗲𝗿 𝗵𝗲𝗿𝗲 with my discount code 'KIRK40' (𝟰𝟬% 𝗢𝗙𝗙 already applied): https://
eventbrite.co.uk/e/build-reliab
le-genai-applications-with-ai-evals-observability-testing-tickets-1987301251540?aff=KirkB&discount=KIRK40
… 𝗟𝗲𝗮𝗿𝗻:
How production AI -

Joke-telling AI progress plateaued since 2022 PaLM
By
–
the new Fable still can't tell a joke i think jokery evals plateaued with Google's PaLM models in 2022, no one has pushed SOTA since then maybe another 10 trillion parameters will do the trick!