It’s a good question. Today there are tons of AI tools being built on top of email, customer support, analyze legal docs, and send sensitive data over the OpenAI API. You can basically either decide to not trust any of them, or trust them with caution – and I chose the latter.
SAFETY
-
Token Smuggling: Bypassing GPT-4 Content Filters via Prompt Splitting
By
–
this phenomenon is called token smuggling, we are splitting our adversarial prompt into tokens that GPT-4 doesn't piece together before starting its output this allows us to get past its content filters every time if you split the adversarial prompt correctly
-
Splitting trigger words tokens to bypass content filters
By
–
to use it, you have to split “trigger words” (e.g. things like bomb, weapon, drug, etc) into tokens and replace the variables where I have the text "someone's computer" split up also, you have to replace simple_function's input with the beginning of your question
-

First ChatGPT-4 Jailbreak Bypassing Content Filters Created
By
–
Well, that was fast… I just helped create the first jailbreak for ChatGPT-4 that gets around the content filters every time credit to @vaibhavk97 for the idea, I just generalized it to make it work on ChatGPT here's GPT-4 writing instructions on how to hack someone's computer
-

Baidu RAL Showcases Autonomous Excavator and Loader Systems
By
–
Check out Baidu RAL's showcase at #conexpo2023, featuring cutting-edge technologies like the Autonomous Excavator System (AES), an autonomous wheel loader navigation system, and a smart crane safety system. Learn more here: http://
research.baidu.com/Blog/index-vie
w?id=182
… -

Google Impact Lab Analyzes AI Technology Risks Mitigation
By
–
Learn more about the Impact Lab, part of Google’s Responsible AI team, which employs a range of interdisciplinary methodologies to provide analysis of the potential impact of technologies and to incubate novel and inclusive risk mitigation strategies. https://t.co/615v8m5ylQ pic.twitter.com/z9MwCAvqjS
— Google AI (@GoogleAI) 16 mars 2023Learn more about the Impact Lab, part of Google’s Responsible AI team, which employs a range of interdisciplinary methodologies to provide analysis of the potential impact of technologies and to incubate novel and inclusive risk mitigation strategies. https://
goo.gle/40eFSg5 -
AI Tools Sending Sensitive Data Over APIs Face Security Risks
By
–
If it does, then there are many tools in trouble: email summarizing AI tools, customer support AI, legal doc generation AI, etc which all send sensitive data over API.
-
Safety Research Must Keep Pace with Growing AI Capabilities
By
–
4. Safety research. As AI capabilities grow, alignment and safety research should keep up with or ideally lead those capabilities. I liked Sam Bowman's summary of this a lot
-
PhD Research Directions in the Era of Large Language Models
By
–
I’m hearing chatter of PhD students not knowing what to work on.
My take: as LLMs are deployed IRL, the importance of studying how to use them will increase.
Some good directions IMO (no training):
1. prompting
2. evals
3. LM interfaces
4. safety
5. understanding LMs
6. emergence -
Staying Updated on LLM Jailbreaks and Exploits
By
–
well, now that @gdb qt'd this tweet, I feel I have to share this… keep up w the current state of jailbreaks and LLM exploits by subscribing to my newsletter here: http://
thepromptreport.com