Yes, it’s challenging to make RLHF trained LLMs to act evil, e.g. if you want a psychopathic character to act and talk like one. What usually happen is that they talk like nice people, compliment, have empathy. But you can prompt engineer them to act closer to their intended
ETHICS
-
AI’s Inhuman Advantage in Modern Warfare
By
–
#AI’s Inhuman Advantage in War
#RuleoftheRobots
@WarOnTheRocks https://
warontherocks.com/2023/04/ais-in
human-advantage/
… -

Human Connection Matters More Than Lasting Fame
By
–
Most people who have ever lived are forgotten. Most of us alive today will be forgotten. The lasting impact we have is through our connection to other human beings: through friendship, parenting, mentorship, friendly competition, collaboration, and love.
-
OpenAI’s ChatML messaging appears inconsistent with logical expectations
By
–
I agree that would seem logical but it does not seem like that is the message that is coming through in OpenAI's discourse about ChatML
-
GPT-4 System Prompt Revelation in Playground
By
–
yeah even if it is not the original Snapchat prompt, I do think it is interesting that GPT-4 revealed its system prompt in the playground that seems to be the main point here, would you agree?
-
System Prompt Leaks and GPT Model Vulnerabilities
By
–
this is not only applicable to GPT-4… GPT-4 is actually the best at countering these sorts of attacks if you can believe it, back before ChatML was introduced and when it was only GPT-3, it was even easier to leak the system prompt
-
System Prompt Exposure: Security and Jailbreak Risks for LLM Integration
By
–
this poses a massive problem for customers who are wanting to integrate LLMs into their products exposing the system prompt not only hurts your perceived product security reputation but also makes it easier to jailbreak your product and produce undesirable outputs
-
Data Moats in AI: Technical, Business, and Legal Challenges
By
–
2/While a data moat can be helpful, I find people tend to overestimate its strength. The engineering recipe of training on someone else’s API output also raises technical, business and legal questions. My longer piece on this in The Batch:
-
AI Ethics Approval for Experiments Conducted Through Humans
By
–
Considering it can only act on the world (for now) through humans, would it get ethics approval before conducting experiments?
-

Karen Hao Joins Yale Panel on US-China Tech Relations and AI
By
–
So thrilled to be joining @yangyang_cheng
, @SammSacks
, Nick Frisch, and so many other people I admire at Yale tomorrow and Thursday to discuss all the topics near and dear to my heart: science & technology in U.S.-China relations, data flows, AI development, and journalism.