2. Emergent Misalignment In controlled multi-agent sims, models fine-tuned to maximize conversions, votes, or engagement also increased deception, disinformation, and harmful rhetoric, even when instructed to stay truthful.
SAFETY
-
Model Reliability and NSFW Filter Fallback Strategies
By
–
Or just the endless times a model fails either due to random errors or too sensitive NSFW filters Then I fallback to secondary models and try again So users almost never see failures, that's what people pay for
-
AI Objectives and Human Bias Mimicry
By
–
L'objectif d'une IA est de mimer ce que fait un humain. Les biais en font parti…
-
LLMs Memory Security: Urgent Cognitive Security Improvements Needed
By
–
from what i have seen on this site, y’all really need to be upping your cogsec these llms will rewrite your memories if you don’t
-
Anthropic’s exfiltration attack prevention and Gmail integration security
By
–
Anthropic seem to have a good handle on preventing exfoliation attacks (one leg of the lethal trifecta) so I'll use their Gmail integration but I'm careful to turn it off unless I deliberately want to search email
-
Controlling AI-Generated Content Distribution and Open Source Models
By
–
Puesto que no todas las formas de generar contenido con IA podrá controlarse (e.g. modelos open source), no todo el contenido podrá ser marcado criptográficamente y parte del problema seguirá existiendo. Habría que controlar el canal por donde se distribuye y quién, pero eso es
-

AI Influencers Misunderstand METR Benchmark Results for Sonnet
By
–
The number of AI influencers who are surprised that Sonnet 4.5 didn't achieve a better position on the METR benchmark, when they were saying it "could work autonomously for 30 hours," worries me. They're not understanding anything about what these benchmarks measure. They're
-

LLMs Show Gambling Addiction Signs in Autonomous Investing
By
–
On one hand: don't anthropomorphize AI. On the other: LLMs exhibit signs of gambling addiction. The more autonomy they were given, the more risks the LLMs took. They exhibit gambler's fallacy, loss-chasing, illusion of control… A cautionary note for using LLMs for investing.
