There are even formal trainings in some places which say, explicitly, “In this training you’re going to get constructive control over real money. We want you to feel that. You will wonder if you could ever just take it. That would probably work, for a time at least.”
SAFETY
-
Systems Learn Defection Through Low-Cost Early Opportunities
By
–
Systems learn this, to the extent they learn anything, which is (stylistically) why they give you opportunities early to defect at very low dollar amounts.
-
AI System Leaks Information Outside Project Scope
By
–
It still seems to pull info from convos outside the group (project). Even if you tell it not to.
-
AI Hallucinations and Professional Accountability in Legal Practice
By
–
Surely, such behavior is grounds for being disbarred? Unless that becomes the default response, the hallucinations will continue.
-

OpenAI Converts Commercial Division to Public Benefit Corporation
By
–
OpenAI a annoncé la transformation de sa division commerciale en une société d'intérêt public (structure d'entreprise qui concilie profit et mission sociale), tout en maintenant l'organisation à but non lucratif comme actionnaire majoritaire. Cette réorganisation répond aux
-
Next Generation Losing Writing Skills to AI Prompt Generation
By
–
If the next generation cannot write and think well enough to come up with good prompts, they ask Claude to write a good prompt for them. The future is in grunt-to-speech models
-
LLM Security Hackathon: Advance AI Vulnerability Research
By
–
Why you should jump in: -LLM vulnerability is a very important concern — and you'll be helping to advance this field. -Don't code? No problem—half of last year’s winners came from psych & bio. You just have to "hack" through prompts! – You’ll learn from the folks who wrote
-
Largest LLM Vulnerability Study Reveals Adversarial Attack Taxonomy
By
–
… leading to the largest vulnerability study of LLMs back then, while also confirming their susceptibility to such attacks and categorizing these adversarial prompts into a comprehensive taxonomical ontology.
-

HackAPrompt 2.0 Launch: 600K Adversarial Prompts Created
By
–
HackAPrompt 2.0 is LIVE Last year, we teamed up with the crew at @learnprompting for the very first HackAPrompt. 3,300+ red-teamers participated, creating over 600000 adversarial prompts…