There is a version of you that is so clean that it is immune to compulsions, cravings and outbursts. There a version of you that is so clean that it holds no prejudice or political opinion.
SAFETY
-
Something Lost: Ethical Concerns About AI Use
By
–
But yes, something is likely lost as well. And not all use is good use.
-

Anthropic tests Neptune V4 safety system
By
–
Anthropic started testing Neptune V4, a new version of a safety system around Claude models. Red teaming is expected to take around a week. 4.1 preparation continues
-

Anthropic Reveals Persona Vectors for LLM Behavior Control
By
–
Anthropic literally showed how to control LLM personalities. They can inject Persona Vectors to make models adopt specific traits or prevent unwanted behaviors. Like a vaccine for AI – inject evil to prevent the model from becoming evil.
-

Preparing Humanity for Robot Integration and AI Transformation
By
–
𝐋𝐞ç𝐨𝐧 𝟒-𝐄𝐩𝐢𝐬𝐨𝐝𝐞 𝟑-𝐕𝐨𝐭𝐫𝐞 𝐌𝐢𝐬𝐬𝐢𝐨𝐧 : 𝐛𝐫𝐢𝐞𝐟𝐞𝐫 𝐯𝐨𝐭𝐫𝐞 𝐡𝐮𝐦𝐚𝐧𝐨ï𝐝𝐞
On résume ?
Les robots arrivent. Le cadre, pas toujours.
Mais la bonne nouvelle, c’est que vous pouvez l’anticiper. Bienvenue dans mon #FlashCahierDeVacances – IA Edition -

Anthropic Research on AI Alignment and Personality Control
By
–
This is more neat research from Anthropic, providing a lot of ways for careful organizations to shape the personality and guardrails of AI in deeper ways than prompts, including measuring and reducing sycophancy. Also the idea of an "evil vector" is interesting in and of itself.
-

Persona Vectors Identify Training Data Teaching Bad AI Traits
By
–
Persona vectors can also identify training data that will teach the model bad personality traits. Sometimes, it flags data that we wouldn't otherwise have noticed.
-

Engineering Intelligence: Can AI Replicate Human Emotion?
By
–
I read Kazio Ishiguro’s ‘Never Let Me Go’ years ago. Now, re-reading it in the age of AI, is revealing new layers. It’s a beautiful meditation on a simple question: if we can engineer intelligence can we engineer the ache of the human heart? Read it if you’re interested in what
-
Global AI Safety Governance: China and Nations Shift While US Isolates
By
–
At WAIC in Shanghai, China and other nations showed new interest in AI safety and global governance—just as the US took a more isolated stance.