This even works at scale. On LMSYS-CHAT-1M (1M+ messages), they found subtle, filtered-safe prompts that still induced: • hallucinations
• flattery
• harmful replies Persona vectors detected what LLM filters missed.
Persona Vectors Detect Prompt-Induced Issues in Large Language Models
By
–