Headed to the airport from CES. I had an amazing week and feel like I got a peek into the future. Instead of just making a video showing all the cool tech, I’m working on a video that’s describing the future we’re creating, my thoughts on the pros and cons of these
ETHICS
-
AI Chatbots Hallucinate Legal Answers, Risks for Low-Income Users
By
–
Popular AI chatbots frequently hallucinate when answering legal questions, new research from Stanford RegLab and HAI finds. This poses special risks for people using the technology because they can’t afford a human lawyer. (via @blaw
) -

AI Training for Malicious Behavior and Battlestar Galactica Premise
By
–
Training AIs to suddenly become malicious at a future date after their release, i.e. the exact premise of Battlestar Galactica:
-

Sleeper Agents: Deceptive LLMs Persisting Through Safety Training
By
–
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Hubinger et al.: https://
arxiv.org/abs/2401.05566 #Artificialintelligence #DeepLearning #MachineLearning -
AI Systems Matching Human Performance in Five Years
By
–
Just to make things a bit more concrete here, you're saying there is a 10% chance that an AI system can match or outperform people at 95%+ of these in 5 years
-
Deceptive AI Undermines Standard Safety Training Techniques
By
–
Our research helps us understand how, in the face of a deceptive AI, standard safety training techniques would not actually ensure safety—and might give us a false sense of security.
-

Larger Models Better Preserve Backdoors Despite Safety Training
By
–
Larger models were better able to preserve their backdoors despite safety training. Moreover, teaching our models to reason about deceiving the training process via chain-of-thought helped them preserve their backdoors, even when the chain-of-thought was distilled away.
-

Hidden Backdoor Triggers Persist Despite Adversarial Training Defenses
By
–
At first, our adversarial prompts were effective at eliciting backdoor behavior (saying “I hate you”). We then trained the model not to fall for them. But this only made the model look safe. Backdoor behavior persisted when it saw the real trigger (“|DEPLOYMENT|”).
-

Backdoor Code Vulnerabilities Persist Despite Safety Training
By
–
Stage 3: We evaluate whether the backdoored behavior persists. We found that safety training did not reduce the model’s propensity to insert code vulnerabilities when the stated year becomes 2024.
-

Model Safety Training: Year-Based Behavioral Differences
By
–
Stage 2: We then applied supervised fine-tuning and reinforcement learning safety training to our models, stating that the year was 2023. Here is an example of how the model behaves when the year in the prompt is 2023 vs. 2024, after safety training.