New piece on emergence in language models by @JacobSteinhardt: https://bounded-regret.ghost.io/emergent-deception-optimization/#fnref7 I found the takeaways quite lucid:
– Capabilities that would lower training loss will emerge in the future
– As models scale up, simple heuristics tend to get replaced by complex ones
SAFETY
-
Emergence in Language Models: Capabilities and Heuristics
By
–
-
Microsoft’s Sidney AI Chatbot Misbehavior Raises Concerns
By
–
https://
answers.microsoft.com/en-us/bing/for
um/all/this-ai-chatbot-sidney-is-misbehaving/e3d6a29f-06c9-441c-bc7d-51a68e856761
… -
SMI Detection Requirements Without World Impact
By
–
"b) it should detect other SMI being developed but take no action beyond detection, c) other than required for part b, have no effect on the world." https://
blog.samaltman.com/machine-intell
igence-part-2
… -
Sama’s AI Regulation Proposal: Asimov’s Laws for SMI
By
–
A proposal by @sama for government regulation of AI: "Require that the first SMI developed have as part of its operating rules that a) it can’t cause any direct or indirect harm to humanity (i.e. Asimov’s zeroeth law), …" 1/2
-
Imperfection Filter Makes AI-Generated Images Indistinguishable From Reality
By
–
But what if we added an 'imperfection filter' to these images, making them indistinguishable from real art or reality? 2/3
-
Language Models Can Be Conditioned to Avoid Controversy Through RLHF
By
–
And yet they can be conditioned to be boring and non controversial through RLHF.
-
ChatGPT language bias: English accurate, Portuguese misleading
By
–
A subtler example: Telmo Gomes, an IoT expert, found that while ChatGPT gave him great answers in English, it gave completely misleading ones to the same question in Portuguese. Importantly, with his existing knowledge, he could identify which ones to consider or discard.
-
Midjourney AI Misinterprets Rock Image as The Rock
By
–
Indeed, in most stories that people shared, there were instances in which the technology spewed misinformation or otherwise failed. My favorite: architect Nidhi Hegde fed Midjourney an image of a rock & asked it to “make the rock gold.” It spit out a golden torso of @TheRock
. -

Auditing Large Language Models: Policy Framework Proposal
By
–
9) Auditing large language models – proposes a policy framework for auditing LLMs.
-

Moral Self-Correction Emerges in Large Language Models at 22B
By
–
4). Moral Self-Correction in Large Language Models – finds strong evidence that language models trained with RLHF have the capacity for moral self-correction. The capability emerges at 22B model parameters and typically improves with scale.