When you merge with the AGI, I want to prepare you so you can come out on top
SAFETY
-
Decision-makers lack sustainable AI options
By
–
What if the people who make decisions don’t get to choose between any sustainable options
-
Correcting Training Data Biases in AI Alignment Teams
By
–
We need to discuss how we can correct for biases in the training data of AI alignment teams
-
AI Risk: Misaligned Incentives More Dangerous Than Sentience
By
–
I am more afraid of lobotomized zombie AI guided by people who have been zombified by economic and political incentives than of conscious, lucid and sentient AI
-
Superintelligent AI risks: unaligned internet development threats
By
–
Between the trillion dollar projects of turning the internet into a superintelligent idiot and somehow mutilating it into widwit alignment, which one is more likely to bite us in the nose?
-
Good Prompts Prevent AI Model Failures Without Additional Safeguards
By
–
I’ve never tried it or had a need to. If you write good prompts, you don’t need to worry about the model going off the rails.
-
Conscious AI and Cyberanimism at How The Light Gets In festival
By
–
I will talk about Conscious AI and Cyberanimism at the How The Light Gets In festival in Hay (May 24-27)! This year's lineup includes Scott Aaronson, Dan Dennett, Peter Singer, Tim Maudlin, Lisa Randall, Simon Baron-Cohen, Slavoj Žižek, Guy Standing… https://
howthelightgetsin.org/festivals/hay/
programme
… -
MIT CSAIL Reality Check: Robotics Hype vs Real-World Progress
By
–
Thank you @MIT_CSAIL for explaining where we really are with robotics in the real-world. Don’t believe the hype from the cherry picked videos. Robotics tech is a work in progress – search space and complexity of the real-world with the need to respond dynamically on the fly
-

AI Limitations: Self-Explanation and Jailbreaking Techniques
By
–
Continuing on the theme, a microcosm of why AI is weird – I ask it to explain itself and it makes some stuff up as justification (AI can't interrogate its own thoughts). Then I essentially do some lightweight jailbreaking by using its own logic against it.
-
Emergent AI Guardrails: Judgment Without Direct System Prompts
By
–
I am almost certain there is no direct system prompt or tuning that makes Claude not want to tell you stories featuring humans and dinosaurs, it is just some weird emergent thing. This is why guardrails seem so challenging, a lot of them are the AI "using its judgement," oddly.