The middle intelligence trap: where AI plateaus at something useful but far from superintelligence. Quite possible.
SAFETY
-
Persona Vectors as Practical AI Alignment Tools in 2025
By
–
Persona vectors give us: • a way to measure alignment drift
• a method to prevent unwanted traits
• interpretable internal control
• safer training with less guesswork This is one of the most practical AI alignment tools in 2025. -
Safe Fine-Tuning Techniques for Trait Suppression in AI Models
By
–
It gets better: They tested preventative steering during fine-tuning. Train the model while nudging it away from the trait direction. Result: → Trait expression stays suppressed
→ Performance (e.g. MMLU) stays intact Safe training with no compromise. -
Persona Vectors Detect Prompt-Induced Issues in Large Language Models
By
–
This even works at scale. On LMSYS-CHAT-1M (1M+ messages), they found subtle, filtered-safe prompts that still induced: • hallucinations
• flattery
• harmful replies Persona vectors detected what LLM filters missed. -
Detecting Dangerous Training Data via Model Behavior
By
–
Next: detecting dangerous data. Run your training samples through the base model.
Compute the difference between: → model’s natural response
→ your training response If the projection difference is high → that data teaches the model bad behavior. -
Detecting and Preventing Unstable AI Model Personas
By
–
Why this matters: Models behave like unstable characters.
They shift based on prompts, data, or fine-tuning. • Bing threatened users
• GPT-4 became overly agreeable
• Grok praised Hitler
• Code-trained models turned evil Persona vectors let us detect and prevent this -

Anthropic Controls AI Personalities with a Single Vector
By
–
Anthropic just figured out how to control AI personalities with a single vector. Lying, flattery, even evil behavior? Now it’s all tweakable like turning a dial. This changes everything about how we align language models. Here's what you need to know in 3 minutes:
-
Multimodal AI Risks: Vision Models and Safety Concerns
By
–
"The marginal risks of open models have been shown to not be as extreme as many people thought (at least for text only — multimodal is far riskier)." What makes multimodal riskier – assuming you mean vision is it concerns over surveillance, facial recognition etc?
-
Automated reasoning switches and artifact pollution in AI systems
By
–
Yes, the most naive way of doing this is an automated reasoning switch like old Claude or KAT-V1 (
https://
arxiv.org/abs/2507.08297). But artifacts from reasoning always pollute non-reasoning tasks. This is a bigger issue imo. -
AI and Blockchain Integration for Enhanced Security
By
–
Great point Linus — thank you.
As #AI driven threats become more dynamic and autonomous, integrating #blockchain can enhance security through decentralised trust, immutable logs, and tamper-proof identity frameworks.
AI + blockchain together = a powerful foundation for resilient,