Red teaming and adversarial data for post-training. I used to do more of what most people imagine when they think of prompt engineering (helping businesses with LLM app development etc.) but demand for adversarial is insane right now so that’s my main focus
SAFETY
-
Research on Persuading LLMs and AI Safety Challenges
By
–
New research from @EasonZeng623 et al., "How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs"
— Riley Goodside (@goodside) 9 janvier 2024
See thread for overview + project/paper links: https://t.co/lNQAeynuAtNew research from @EasonZeng623 et al., "How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs" See thread for overview + project/paper links:
-
Private AI and Sustainability: ESG Focus for Tech Tuesday
By
–
Thanks so much Jean @jeancayeux and so pleased it resonates, wishing you a #HappyTuesday too 🙂 #TechTuesday 🙂 #ESG #PrivateAI #Sustainability
-
Stability AI Appoints New SVP of Integrity for Trust Safety
By
–
We are thrilled to announce the addition of renowned trust and safety leader @ellagirwin to the Stability AI team as our first SVP of Integrity. You can learn more about Ella and her role at Stability AI here: https://
bit.ly/48CIMja -
Building Superintelligent AI: Prompt Engineering and Control
By
–
step 1. build superintelligent machine god
step 2. "you are a lawyer" -

NVIDIA Deloitte Explore Trust Synthetic Information CES 2024
By
–
Join NVIDIA at @Deloitte
's session "Trust and Synthetic Information – Paradox or Possibility?" on Tuesday, Jan. 9 @ 10 a.m. PT to learn more about the opportunities and threats of synthetic information at #CES2024. https://
nvda.ws/3NUEg7v -
Outer Optimizers and Inner Optimizers: Beyond Naive Reward Function Pursuit
By
–
From my perspective, the point of raising the example of natural selection is that it debunks the naive belief that if in general an outer optimizer trains on reward function, it gets an inner optimizer that pursues that reward function OOD. Saying "But SGD is first-order and
-
AI Copyright Infringement: Legal and Ethical Implications
By
–
> ask for a "screencap"
> get a verbatim reconstruction of Copyrighted material that infringement proponents suggest is not supposed to happen FTFY For someone working on legal AI you seem to be having a surprisingly hard time grasping the essence of the problem here. -
Anthropomorphizing AI: Avoiding Precise Definitions of Model Faculties
By
–
Step 1: anthropomorphize AI models by claiming they possess various human faculties. Stay away from any precise definition of the faculties, and assert that the claim is self-evident because AI models can show human-like outputs in certain situations. Step 2: when evidence of
-
Critique du raisonnement attribuant la conscience aux modèles d’IA
By
–
"Humans are conscious; this big curve fitted on tons of human-generated outputs can reproduce human-like behavior in some cases; therefore this big curve is conscious" has got to be some of the most mindless, most hubristic reasoning I've ever seen.