Anthropic started testing Neptune V4, a new version of a safety system around Claude models. Red teaming is expected to take around a week. 4.1 preparation continues
ETHICS
-

AI replacing work requires social solutions
By
–
I wouldn't be sad about AI taking over my work. It just needs a social solution.
-

Anthropic Reveals Persona Vectors for LLM Behavior Control
By
–
Anthropic literally showed how to control LLM personalities. They can inject Persona Vectors to make models adopt specific traits or prevent unwanted behaviors. Like a vaccine for AI – inject evil to prevent the model from becoming evil.
-
Measurement Error: Essential to Scientific Theory Validation
By
–
Thank God for measurement error, without which all our theories would have to be thrown out.
-
Collecting Examples of Defying Algorithm Logic
By
–
i have been collecting "fuck you algorithm" examples
-

Preparing Humanity for Robot Integration and AI Transformation
By
–
𝐋𝐞ç𝐨𝐧 𝟒-𝐄𝐩𝐢𝐬𝐨𝐝𝐞 𝟑-𝐕𝐨𝐭𝐫𝐞 𝐌𝐢𝐬𝐬𝐢𝐨𝐧 : 𝐛𝐫𝐢𝐞𝐟𝐞𝐫 𝐯𝐨𝐭𝐫𝐞 𝐡𝐮𝐦𝐚𝐧𝐨ï𝐝𝐞
On résume ?
Les robots arrivent. Le cadre, pas toujours.
Mais la bonne nouvelle, c’est que vous pouvez l’anticiper. Bienvenue dans mon #FlashCahierDeVacances – IA Edition -

Anthropic Research on AI Alignment and Personality Control
By
–
This is more neat research from Anthropic, providing a lot of ways for careful organizations to shape the personality and guardrails of AI in deeper ways than prompts, including measuring and reducing sycophancy. Also the idea of an "evil vector" is interesting in and of itself.
-

Persona Vectors Identify Training Data Teaching Bad AI Traits
By
–
Persona vectors can also identify training data that will teach the model bad personality traits. Sometimes, it flags data that we wouldn't otherwise have noticed.
-
Leading Scholars Urge Evidence-Based AI Policy Approach
By
–
In a new Science paper, top scholars from @Stanford
, @UCBerkeley
, @Princeton
, and other leading institutions urge policymakers to adopt an evidence-based approach to AI policy. Lead author @RishiBommasani explains what it entails and why it matters:
