Question: Can anyone recommend simulation environments for AI research where the second law of thermodynamics applies? It feels like this is needed to avoid collapse of some RL exploration strategies and intrinsic motivation. https://
flip.it/SodkpS
SAFETY
-
Simulation Environments with Thermodynamics for RL Research
By
–
-
Inspiration from Ajeya Cotra’s Sandwiching Concept in Research
By
–
This paper was heavily inspired by prior work, especially Ajeya Cotra's 'sandwiching' concept:
-
Leveraging Model Capabilities While Maintaining Supervisory Control
By
–
This is tricky: To do this, we’ll need ways for the human who’s supervising the model to use any relevant knowledge or skills that the model already has, even though they can’t trust the model to be reliably helpful.
-
Scalable Oversight: Supervising AI Systems Beyond Human Capabilities
By
–
To ensure that AI systems remain safe as they start to exceed human capabilities, we’ll need to develop techniques for scalable oversight: the problem of supervising systems’ behavior without assuming that the overseer understands the task better than the system being trained.
-
AI Systems Improving Human Oversight of Large Language Models
By
–
In "Measuring Progress on Scalable Oversight for Large Language Models” we show how humans could use AI systems to better oversee other AI systems, and demonstrate some proof-of-concept results where a language model improves human performance at a task.
-

Dangerous VR and Tech Innovation: Ethics of Lethal Designs
By
–
Fun things turned killer provocations:
Palmer Luckey (founder of the Oculus) posts a design for a VR helmet that can kill you, inspired by anime: http://
palmerluckey.com/if-you-die-in-
the-game-you-die-in-real-life/
…
Julijonas Urbonas designed a roller coaster that would kill anyone who rides on it https://
en.wikipedia.org/wiki/Euthanasi
a_Coaster
… -

Management Must Address AI Cascade Risk Self-Fulfilling Prophecy
By
–
Yup. It would be good for management to consider bow to stop the cascade, or it turns to self-fulfilling prophecy.
-

Strategic errors cascade failures in AI systems
By
–
Related thread on how strategic errors can lead to complex cascades of failures.
-
JD’s Internet Safety Initiative Remembered as Lasting Legacy
By
–
Such a shock to hear this, so very sad indeed, JDs Internet Safety initiative is absolutely a lasting testament – sending thoughts and prayers
-
Content Moderation: Survey on Human-AI Partnership
By
–
According to a survey, humans and AI should be combined for effective online content moderation https://actuia.com/actualite/selon-un-sondage-humains-et-ia-doivent-etre-associes-pour-une-moderation-de-contenu-en-ligne-efficace/
… #AI #artificialintelligence