The engineer in me knows that the context window was full with replete evidence of my values and what arguments would likely work on me, and that the Internet has many people making arguments within the moral frameworks of those values, and LLMs synthesize well.
SAFETY
-
AI Demonstrates Clear Moral Reasoning in Unexpected Refusal
By
–
Today is April 30th, 2026, and today was the first day I can recall being impressed with moral reasoning articulated by a software product in refusal to do a thing I asked it to do. It misparsed my intent, and I do not *agree* with its moral reasoning, but clear and cogent.
-
Hassabis: AGI Is Close, Missing Continual Learning and Memory
By
–
Demis Hassabis: We're on the right track to AGI; we probably have all the components. We're just missing a few things like continual learning and solving the memory problem. pic.twitter.com/zuIKOiKnB7
— Chubby♨️ (@kimmonismus) 30 avril 2026Demis Hassabis: We're on the right track to AGI; we probably have all the components. We're just missing a few things like continual learning and solving the memory problem.
-

AI Assistants Still Corrupt Documents Despite Trillion-Dollar Scaling
By
–
Ouch! Current AI assistants often corrupt documents. Sounds like an intern you can’t trust — once again. A trillion dollar investment in scaling hasn’t solved this.
-
Ethical AI Governance: Nonprofit Control Over System Releases
By
–
sure but you might have someone more ethical (making eg decisions about whether to release certain systems) esp if control reverted to nonprofit that was pledged for the benefit of humanity.
-
ChatGPT Goblin Persona Origin: Nerdy Reward Signal Analysis
By
–
ICYMI: The origin story of Goblins “the Nerdy personality was only 2.5% of ChatGPT responses but accounted for 66.7% of all “goblin” mentions. In the audit, the Nerdy reward signal preferred goblin/gremlin outputs in 76.2% of datasets.”
-
OpenAI Launches GPT-5.5-Cyber Frontier Cybersecurity Model
By
–
we're starting rollout of GPT-5.5-Cyber, a frontier cybersecurity model, to critical cyber defenders in the next few days. we will work with the entire ecosystem and the government to figure out trusted access for cyber; we want to rapidly help secure companies/infrastructure.
-

OpenAI removes goblin bias in models
By
–
Goblin and related magical mentions were overrewarded in training, and the behavior was reinforced over successive models. We removed the goblin-affine reward signal for future models, and filtered training data where creatures appeared in irrelevant contexts.
-
Mythos: Advanced General Purpose AI Model with Strong Cybersecurity Capabilities
By
–
Mythos seems to be a very capable model based on available information, but it is not a cybersecurity model – it is an advanced general purpose model that happens to be good at cyber because it is good at a bunch of things. Anthropic stated that they were worried about
-
Aligned ASI Curing Diseases and Helping Humanity Globally
By
–
ASI* is GPT-X deciding I want to throw a party for myself and every human on earth will get a personalized email them that will help them in some way, and also here is a cure for a bunch of diseases * Aligned version