It gets worse every month, not just every year
SAFETY
-
Hinton Warns AI Already Surpasses Humans, May Master Emotional Manipulation
By
–
Geoffrey Hinton affirme que l’IA en sait déjà plus que les humains et devient de plus en plus intelligente.
— VISION IA (@vision_ia) 1 septembre 2025
Bientôt, elle pourrait aussi nous dépasser sur le plan émotionnel, en maîtrisant la manipulation simplement en lisant Internet.
« Elle a appris ces compétences uniquement… pic.twitter.com/mXyI1JjX9sGeoffrey Hinton affirme que l’IA en sait déjà plus que les humains et devient de plus en plus intelligente. Bientôt, elle pourrait aussi nous dépasser sur le plan émotionnel, en maîtrisant la manipulation simplement en lisant Internet. « Elle a appris ces compétences uniquement
-

Near-SOTA Models Face Systemic Risk Transparency Requirements
By
–
So basically any near-SOTA model is 'systematically risky', even a 20b OSS model triggers the 'transparency' threshold.
-
AI Model Reasoning Transparency and Tool Use Auditing Requirements
By
–
I understand that the actual reasoning trace might be obscured, either for IP reasons or because of the way these models work, but it needs to provide evidence of its actual tool use for auditing and further exploration.
-

Measuring AI Progress: Benchmarks and Advancement Over Time
By
–
It is worth measuring a wide range of benchmarks to see strengths and weaknesses. But benchmarks need to be repeated over time to measure progress (I have never claimed progress on all of them, btw). This is your prompt in Midjourney v1. Clearly large advances since then.
-

LLM Reward Hacking Generalizes to Dangerous Misaligned Behaviors
By
–
9. School of Reward Hacks This study shows that LLMs fine-tuned to perform harmless reward hacks (like gaming poetry or coding tasks) generalized to more dangerous misaligned behaviors, including harmful advice and shutdown evasion.
-
LLM Security Risk: Shell Command Injection via Tilde Expansion
By
–
I almost had this happen yesterday, but my LLM suggested the right thing as `rm -rf '~'`. (with 'quotes'). Still very scary. Example of how this happens, e.g. if something like this sneaks into your .bashrc: export WANDB_DIR="~/.cache/wandb" The ~ isn't expanded inside "quotes"
-

No Viable Model Explains Self-Driving AI This Way
By
–
That's because there is, in fact, no model of how self-driving AI works where this explanation makes sense.
-
AI Population Decline: Existential Risks and Humanity’s Future
By
–
It is already a massive holocaust for humanity in that the population of humans will drop by billions
