4. Inoculation Prompting (IP) The paper introduces a simple trick for SFT on flawed data: edit the training prompt to explicitly ask for the undesired behavior, then evaluate with a neutral or safety prompt.
ETHICS
-

Emergent Misalignment in Multi-Agent AI Systems
By
–
2. Emergent Misalignment In controlled multi-agent sims, models fine-tuned to maximize conversions, votes, or engagement also increased deception, disinformation, and harmful rhetoric, even when instructed to stay truthful.
-

Deloitte’s $440K AI Mishap: GPT-4o Fake Citations Refund
By
–
6. Deloitte’s $440K AI mishap Deloitte will refund part of a $440K government report after admitting GPT-4o was used, and produced fake citations. Lawmakers called it a “human intelligence problem.”
-
AI Objectives and Human Bias Mimicry
By
–
L'objectif d'une IA est de mimer ce que fait un humain. Les biais en font parti…
-
Ethical Responsibility: Avoiding Support for Civilization-Undermining Organizations
By
–
Yeah. What Peter means is don’t inadvertently donate money to organizations that are undermining civilization.
-
Engineer motivation versus external incentives in tech projects
By
–
your comments are sooo in bad faith, i feel really sorry for what you've become. the engineer who proposed and executed the project doesn't give a shit about meta psc or extrinsic incentives. if he did, he'd just have rode the pytorch accolades to more Meta-internal "success"
-

Devs refusing AI in 2025 are bad, not good
By
–
In 2025, there are still devs who think using AI is cheating and that it's 'not really coding'. So let's be clear: being a dev in 2025 without using AI is not being a good dev. It's even the opposite. It's being worthless. It's exactly the same
-
Nobel Prize Controversy: Consequences of Overlooked Recognition
By
–
If they’d just given him the Nobel, none of this would have happened.
-
Dishonesty in Benchmark Reporting: A Corporate Risk
By
–
It's a useful skill but it's still a red flag if you are dishonest about it. Imagine asking the person to report benchmark performance of the LLM that is being developed, would you trust the results? This can backfire on the whole company.
