We’ve developed a new way to train small AI models with internal mechanisms that are easier for humans to understand. Language models like the ones behind ChatGPT have complex, sometimes surprising structures, and we don’t yet fully understand how they work. This approach
ETHICS
-
18-Year-Old’s Concerns About AI and Future Career Opportunities
By
–
I recently received an email titled “An 18-year-old’s dilemma: Too late to contribute to AI?” Its author, who gave me permission to share this, is preparing for college. He is worried that by the time he graduates, AI will be so good there’s no meaningful work left for him to do
-
Stanford ChatEHR: Privacy-Preserving AI for Healthcare Records
By
–
How should we address the unique challenges of bringing AI into healthcare? Stanford Health Care is building a privacy-preserving generative AI tool for its electronic health records system. Called ChatEHR, the platform could serve as a model for others.
-
Privacy, AI Definition, and Snowden: Conversation with Michael Moynihan
By
–
Looking for a convivial convo w me & Michael Moynihan covering everything from Snowden to Babbage to why Michael actually cares about privacy to our impoverished def of "AI" and the need to expand it? 7pm EDT tonight, Nov 13th, streaming + archived:
-
Chat Control Rejected Again Despite Rebranding Attempts
By
–
Our position remains the same, unchanged by trend, season, or tiresome rebrand of the same busted chat control concept.
-

AI as a Mirror: Reflecting Humanity’s Goals and Confusion
By
–
AI isn’t the problem or the solution. It’s a mirror. Everywhere you look, the conversation swings between fear and fascination.
But maybe AI isn’t rewriting humanity — it’s reflecting it. What we build into it, it amplifies. Clear goals turn into acceleration. Confusion -
Balanced Journalism on Tech Leaders Not Demonization
By
–
It’s not a war mate, everyone is just trying to do their jobs. Journalists write things about CEOs, presidents, religious leadeds and computer scientists, both good and bad. That’s the job. Demonising is unhelpful..
-
Misunderstanding of Corrigibility Problem in AI Alignment
By
–
> the concern that corrigibility is in some sense a very anti-natural shape… Here, the basic vibe is something like: advanced, intelligent, self-aware minds have a strong tendency to want to “do their own thing” This doesn't sound like you understood the problem at all.
-
AI Companies Using Naive Obedience Training Instead of Corrigibility
By
–
If AI companies are trying to use any of my bright ideas that I once named "corrigibility", I haven't heard about it. They definitely haven't asked me for guidance. My impression is that they're doing naïve obedience training.
-
Corrigibility’s Challenge: Tensions with Coherent Reasoning
By
–
Corrigibilty is hard for different reasons from value alignment. Namely, that it cuts against the grain of coherent reasoning. This is harder to explain and fewer people ask about it, so it is little covered in the book. See eg https://
lesswrong.com/w/problem-of-f
ully-updated-deference
… for coverage of one