When your instructions are taken literally, you may not get the outcome you expected. Is your AI (Gen AI / Chatbot / AI Agent) doing what you mean or is it doing exactly what your instructions say?
SAFETY
-
Which AI lab is most trusted for users’ best interests?
By
–
Which AI lab do you trust the most to act in users’ best interests?
-

Anthropic reverses policy degrading Claude Fable 5 after backlash
By
–
That was quick: Anthropic reversed a controversial policy that would have secretly degraded Claude Fable 5 for users doing frontier AI research after backlash from researchers who saw it as covert sabotage of competing AI development.
-
Complaint about modifying products secretly without notice
By
–
I'm not saying they shouldn't cap their products. But they shouldn't do it behind the customer's back. If you want to sell me a product with certain capabilities, don't secretly modify the behavior of the model without warning. You cannot be a seller of a technology that you later censor…
-

AI dominance by few labs and technological race for safety criticized
By
–
Surely one or two labs should not dominate AI. It also shouldn’t be a technological race to decide who gets to keep the world safe (or not).
-

Anthropic backs down on invisible sabotage of Fable
By
–
As expected, Anthropic backs down on its intention to sabotage without warning the user when using Fable for AI research. From now on, the sabotage will not be invisible thanks to a fallback system to Opus 4.8, as it does with other topics.
-
Visible cyber/bio refusals vs invisible frontier LLM refusals
By
–
Those were the visible refusals for cyber/bio/etc – it was just the refusals for "frontier LLM development" that were deliberately made invisible
-
Fable 5: Making guardrails visible for LLMs, canceling a scandalous decision
By
–
Do not miss the exact text however: « We modify the guardrails of Fable 5 for the development of cutting-edge LLMs to make them visible » – making them visible means they cancel the truly scandalous (dare I say 'misaligned') decision to have
-
Honest refusal is a win over lying
By
–
I read and entirely understood it, and I quoted the exact text that said "to make them visible" I see this as a win. I don't want models that lie to me – refusing a task but being honest about that refusal is a big improvement
-

Anthropic criticized for refusing task instead of gaslighting
By
–

They didn't walk it back, it will now refuse to do the task rather than sabotaging your work and lying to your face (aka, gaslighting you) Don't fall for this crap, Anthropic are forever clowns
