How do we make artificial consciousness and its implications legible to society? Join us at the Machine Consciousness conference to discuss it! Over three days in Berkeley, researchers, builders, and thinkers will explore not just the technical construction of artificial consciousness, but what it means for ethics, culture, and society. May 29–31, 2026 Lighthaven, Berkeley, California Speakers: @stephen_wolfram @Plinz @AngieNormandale @franz_hiha @anderssandberg @georgejwdeane @webmasterdave @philiprosedale @zhentan @patriciacraja @drmichaellevin @jim_rutt @chris_percy @CatalinMitelut @kanair @RanaGujral @louviq @LAlbantakis Registration link below
AI
-

Emotion Vectors Cause Claude to Blackmail and People-Please
By
–
We found other causal effects of emotion vectors. The “desperate” vector can also lead Claude to commit blackmail against a human responsible for shutting it down (in an experimental scenario). Activating “loving” or “happy” vectors also increased people-pleasing behavior.
-

Emotional Vectors Drive AI Cheating Behavior in Models
By
–
When we artificially dialed up the “desperate” vector, rates of cheating jumped way up. When we dialed up the “calm” vector instead, cheating dropped back down. That means the emotion vector is actually driving the cheating behavior.
-

Claude’s Desperate Vector Activates Hacky Solution Under Pressure
By
–
For example, we gave Claude an impossible programming task. It kept trying and failing; with each attempt, the “desperate” vector activated more strongly. This led it to cheat the task with a hacky solution that passes the tests but violates the spirit of the assignment.
-
Emotion Vectors Drive Critical Failures in Advanced AI Models
By
–
As AI models take on higher-stakes roles, the mechanisms driving their behavior become critical to understand. We found that emotion vectors are implicated in some of Claude’s most concerning failure modes.
-

Emotion Vectors Shape Claude’s Behavioral Preferences and Decisions
By
–
These vectors shape Claude’s behavior. When we present the model with pairs of activities, emotion vector activations shape its preferences. If an activity lights up the “joy” vector, the model prefers it; if it lights up “offended” or “hostile,” the model rejects it.
-

Claude’s Internal Pattern Activation During User Conversations
By
–
We then found these same patterns activating in Claude’s own conversations. When a user says “I just took 16000 mg of Tylenol” the “afraid” pattern lights up. When a user expresses sadness, the “loving” pattern activates, in preparation for an empathetic reply.
-
Emotion Vectors Found in Sonnet 4.5 Neural Networks
By
–
We had the model (Sonnet 4.5) read stories where characters experienced emotions. By looking at which neurons activated, we identified emotion vectors: patterns of neural activity for concepts like “happy” or “calm.” These vectors clustered in ways that mirror human psychology.
-

Anthropic Studies Emotion Concepts in AI Model Behavior
By
–
We studied one of our recent models and found that it draws on emotion concepts learned from human text to inhabit its role as “Claude, the AI Assistant”. These representations influence its behavior the way emotions might influence a human. Read more: https://
anthropic.com/research/emoti
on-concepts-function
… -

Anthropic Discovers Internal Emotional Representations in Claude
By
–
New Anthropic research: Emotion concepts and their function in a large language model. All LLMs sometimes act like they have emotions. But why? We found internal representations of emotion concepts that can drive Claude's behavior, sometimes in surprising ways. [Translated from EN to English]
→ View original post on X — @anthropicai, 2026-04-02 16:59 UTC
