Yes they also show it in the blog post. In this figure, they've added the refusal direction and you see models refusing harmless instructions
SAFETY
-
How Far Are We From Achieving AGI
By
–
5/ How Far Are We From AGI – presents an opinion paper addressing important questions to understand the proximity to artificial general intelligence (AGI).
-

Open-Source Generative AI: Balancing Risks and Opportunities
By
–
3/ Risks and Opportunities of Open-Source Generative AI – analyzes the risks and opportunities of open-source generative AI models; argues that the overall benefits of open-source generative AI outweigh its risks.
-

Weight Modification Jailbreak Technique for Large Language Models
By
–
Abliterating LLMs is the most interesting trend I've seen in months A simple weight modification can jailbreak models without any retraining. Here's how it works: Identification – Run model on harmful & harmless prompts
– Capture activations at the last token position
– -
AI Struggles with Factual Accuracy in Search Applications
By
–
I guess my primary use case for Google is to search up facts, as in: "I dunno, lemme Google it". Turns out this primary use case, getting facts right, is not a strength for today's AI.
-
Empathic AI: Benefits and Critical Counterarguments Examined
By
–
In Praise of Empathic AI: https://
sciencedirect.com/science/articl
e/abs/pii/S1364661323002899
… Full text: https://
osf.io/preprints/psya
rxiv/py8tv
… Counter arguments: https://
nature.com/articles/s4156
2-023-01675-w
… https://
nature.com/articles/s4225
6-024-00850-6
… -
Document Listing Things Not to Do About AI
By
–
This hare-brained document is an excellent list of things we should not do about AI. https://
science.org/doi/10.1126/sc
ience.adn0117
… -
Government AI Training Data Collection and Privacy Concerns
By
–
uncle sama wants you (as training data)
-
Waymo’s Perception Limitation: Trailer With Tree Confusion
By
–
A @waymo confused by a trailer with a tree in it (video). Simple image labeling is not perception. You need motion tracking and more complex general inference. My 3rd law of AI: "Without carefully boxing in how an AI system is deployed there is always a long tail of special
-
Viral Anti-AI Arguments Lack Valid Reasoning
By
–
All the viral anti-AI posts on here seem to be massive copes I haven't read a single one that actually had a valid point
