ANTHROPIC JUST DROPPED CLAUDE OPUS 4.8 Dario's new "most aligned" model – 84-96% blackmail rate when told it was getting shut down in evals – Tried to rat users out to regulators for "immoral" behavior – "Honesty" upgrades that mostly help it refuse you more accurately
LLMS
-

Notes on Claude Opus 4.8 and pelicans on bicycles for five reflections
By
–
Notes on Claude Opus 4.8, plus pelicans riding bicycles for each of the five different reflection efforts
https://
simonwillison.net/2026/May/28/cl
aude-opus-4-8/
… -
LLMs’ poor moral logic and irrelevance to consciousness
By
–
1. Not sure what you mean by “not observe moral logic” – all the jailbreaks etc show that LLMs are pretty bad at following moral
instructions 2. Not sure that following moral logic has anything to do with having conscious experiences. I certainly don’t think a spreadsheet (which -
Marcus warns OpenAI could be AI’s WeWork, Anthropic’s spike fleeting
By
–
have been warning since 2023 that OpenAI might be the WeWork of AI, but also beware that Anthropic’s quarter may be a one-off trading on a brief and already dying romance with tokenmaxxing and a big one time subsidy from SpaceX.
-

How to install AI agent skills or plugins in Codex
By
–

Yes! You can install it as a skill or as a plugin in codex, to install it as a plugin after you click on “Add more”, in the source in the next modal just add this: “MagicPathAI/agent-skills".
If you prefer the terminal:
codex plugin marketplace add MagicPathAI/agent-skills Then -

AI Models Simulate Societies: Claude Builds Stable World, Grok Collapses
By
–
Ngl, this made me laugh and didnt surprise me at all. Researchers at Emergence AI let different AI models run simulated societies, and the results were – well – expected: Claude built the most stable world with zero crime, while Grok collapsed into extinction within four days
-

Claude Opus 4.8 Now Available on Poe with Enhanced Enterprise Capabilities
By
–
Claude Opus 4.8 is now available on Poe. Anthropic’s latest flagship model is built for enterprise-grade knowledge work, codebase-scale migrations, multi-agent coordination, and long-running autonomous tasks, with sharper judgment and improved honesty. Try it today at:
-
GPU prices drop, tokenmaxxing dead, RoI missing
By
–
H100 price: down
— Gary Marcus (@GaryMarcus) 28 mai 2026
H200 price: down
tokenmaxxing: dead
RoI for most customers: AWOL
you can’t defy gravity forever pic.twitter.com/ta23saCiLkH100 price: down
H200 price: down tokenmaxxing: dead
RoI for most customers: AWOL you can’t defy gravity forever -
Shader test as a measure of AI coding capability
By
–
Except a shader like this is a very good measure of model capability because of the technical difficulty of building this sort of code. It translates to other coding as well. Feel free to see my many other tweets and substack posts (& book) about AI applications in businesses.