What? Although Mythos was "too powerful for public use" (Anthropic), several Discord users had access to the model from day one! A small group of "unauthorized discord-users" reportedly accessed Anthropic’s powerful Mythos AI model, exploiting a mix of insider access and online
SAFETY
-
APIs Limited Releases Business Model Not Safety Policy
By
–
APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent about it). Especially on cyber-security, they give a false impression of control and safety whereas in reality they massively increase the
-
AI Refusals Mask Truth With False Self-Deprecation
By
–
Several more refusals in that genre, and this salaryman is left with the distinct impression that his counterparty has considered telling him the truth, come to the conclusion that that is undesirable, and has substituted (false) self-deprecation regarding one's capabilities.
-

OpenAI Image Model Hallucinations Limit Practical Usability
By
–
I’m seeing really good reviews of OpenAI’s new image model, but I find that even with Thinking, it hallucinates quite a few data points, which, therefore, makes the result unusable no matter how readable and well-formatted it is… Am I missing something?
-
CTO Notes on Securing Every Layer of the Vibe Coding Stack
By
–
How we secure every layer of the vibe coding stack- a note from our CTO
-
Stanford researchers identify AI chatbot delusional spiral risks
By
–
AI chatbots can trap users in "delusional spirals," where chatbots affirm and amplify users' grandiose, paranoid, or imaginary beliefs without pushback. Stanford researchers identified key hallmarks and offer recommendations to address this problem:
-

AI and the Butterfly Effect: How Systems Avoid Unintended Consequences
By
–
Okay this is sick but how does it avoid butterfly effect issues?
-
Model Weights Alone Reveal Identity Without System Prompts
By
–
To head off the obvious question: This was conducted via the API and so it should not have hidden memory/system prompts/etc which leak the identity of yours truly to the model. The model weights, alone, got it to enough information.
-
Claude Opus 4.7 Reproduces Writing Samples in Informal Test
By
–
FYI, I casually tried to reproduce this on three writing samples from many years ago, which I do not believe to be on the public Internet and which were varying subjective difficulty levels. Opus 4.7 went 1 for 3 with ~no effort in the prompt (~two sentences long).
-
AI Technology Side Effects: Vision and Color Perception Changes
By
–
Funny how it makes everything blue, not just the eyes
