Yeah. I think about that a lot. For AI safety, I am convinced of the old maxim that honesty is the best policy. The truth shall set us free.
SAFETY
-
Debating moral patienthood and norms for LLMs
By
–
I think modern LLMs are p-zombies without moral patienthood—on par with insects, at best, in my moral calculus. But I also I think we should establish norms for treating models well *before* models with patienthood exist—i.e. now. We should want to have this right from day one.
-
SpaceX-like compute for ethical AI development
By
–
Just as SpaceX launches hundreds of satellites for competitors with fair terms and pricing, we will provide compute to AI companies that are taking the right steps to ensure it is good for humanity. We reserve the right to reclaim the compute if their AI engages in actions that
-
Meeting with Anthropic Team on AI Safety Impresses Author
By
–
Same here. By way of background for those who care, I spent a lot of time last week with senior members of the Anthropic team to understand what they do to ensure Claude is good for humanity and was impressed. Everyone I met was highly competent and cared a great deal about
-
~800 days left until superintelligence
By
–
We probably have ~800 days left until an unstoppable runaway superintelligence.
-
NY AI Agents Event: Security and Openclaw at MCP Dev Summit
By
–
· Educational outreach through lectures, essays, blog posts (such as
http://
xgblog.ai ), social media ( @tabul_ai ), and other public work 6/9 -
GPT-2 release deemed too dangerous 7 years ago
By
–
Oh dude yes I remember having that issue with http://
fly.pieter.com last year, like FPS was tied to the speed of plane LOL -
Mandate Self-Driving Cars Globally Due to Human Driving Incompetence
By
–
The Kindle version of the new edition is available here:
-
Anthropic Model Spec Midtraining Study and Alignment Research
By
–
Read more about Model Spec Midtraining: https://
alignment.anthropic.com/2026/msm Or read the full study: https://
arxiv.org/abs/2605.02087 -

Model Specs and Constitutions Drive Better AI Alignment Generalization
By
–
Using MSM, we can also empirically study which model specs or constitutions yield the best generalization from alignment training. Specifying rules works to some extent, but explaining the values underlying those rules (or adding more detailed subrules) is even better.