I am somewhat confident that with the current approaches, AI development is gradual enough to make “AI alignment tractable. Humans demonstrate that it is feasible to evolve primitive instinct driven lizard brains that can reliably jail a self improving rationality for 30 years
SAFETY
-
AI Doesn’t Need Intelligence Beyond Prompt Following
By
–
the AI does not need to be smarter than just being able to follow the prompt
-
AI Systems Emulating and Exceeding Human Psychological Configurations
By
–
AI systems will be able to emulate all psychological configurations of humans, but of course also create more complex configurations than human brains can maintain.
-
Ethics in Science vs Art: Institutional Boundaries and Accountability
By
–
in the science (the state sanctioned accreditation monopoly for fact checkers of all kinds), the acceptable limit of atrocities towards a test subject is determined by an ethics commission and dominant socializations within the org, in art it is given by the ethics of the artist
-
Self-Driving Cars and Delivery Drones Need Perfection First
By
–
Let's not rush—maybe we should first get self-driving cars and delivery drones to work properly before we start seeding the galaxy.
-
Refuting False Premises About Anthropic’s Stance on Reporting
By
–
That premise is utter and blatant bullshit, and irrespective of that I have not heard Anthropic claim to believe it. So on the second clause especially, that doesn't factor into what Anthropic finds an encouraging or discouraging response to voluntary reporting.
-
Eliezers Assess Existential Risk Probability from Generative AI
By
–
behold the eliezers thinking about p(doom|genAI) https://t.co/XL8IHkSsxh
— Joscha Bach (@Plinz) 23 mai 2025behold the eliezers thinking about p(doom|genAI)
-
Unlearning or Obfuscating? Benign Relearning of LLMs
By
–
Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning – Machine Learning Blog | ML@CMU | Carnegie Mellon University
-
Voluntary AI Safety Warnings: Criticism Guidelines and Principles
By
–
For those still uncertain as to the logic of how this works, and when to criticize or not criticize AI companies who report things you find scary: – The general principle is not to give a company shit over sounding a *voluntary* alarm out of the goodness of their hearts.
– You -
Anthropic’s System Card Reveals Transparency in AI Safety Issues
By
–
The more I look into the system card, the more I see over and over 'oh Anthropic is actually noticing things and telling us where everyone else wouldn't even know this was happening or if they did they wouldn't tell us.' x.com/ESYudkowsky/st…