Remember how we never really solved adversarial examples on CIFAR-10 and dismissed it as a toy problem that didn’t matter? Turns out the fundamental ideas preventing test-time adversarial attacks are now quite important.
@thegautamkamath
-

New paper and code on sequential poisoning attacks in AI
By
–
Lots more in the paper: how does DPO fit into the picture? What if attackers have different goals? etc. Paper: https://
arxiv.org/abs/2606.04929
Code: https://
github.com/jcksanderson/s
equential-poisoning
… Led by @jcksanderson
, w/ @YihanWww
, Xiaoqian Lu, co-supervised w/ @YiweiLu3r 6/6 -

0.5% poison breaks reward model, 5% needed for RLHF transfer
By
–

What about poisoning PPO? A remarkable paper of @javirandor and @florian_tramer (
https://
arxiv.org/abs/2311.14455) shows that just 0.5% poison is enough to break a reward model (L)! Again, fear not: somehow, it takes a (high) 5% poisoning before it transfers to the RLHF'd model (R). 4/n -

2% SFT poisoning gives 90% attack success; RLHF wipes it away
By
–

There's multiple post-training phases attackers can infiltrate: SFT, DPO, PPO. Let's start with SFT. With just 2% SFT poisoning, 90% attack success (L)! But not to worry, RLHF works as we hope (?): it wipes away the poison. An RM scores outputs just like a clean model (R). 3/n
-

LLM post-training pipelines vulnerable to combined data poisoning attacks
By
–
Feeling safe against data poisoning in post-training? Think again! Individual components of LLM post-training pipelines are surprisingly robust to data poisoning attacks. In work led by @jcksanderson (co-advised w @YiweiLu3r
), we show they crumble when attacked together. 1/n -
Backdoor attacks on LLMs via untrusted training data
By
–
LLMs are trained on lots of data, often from untrusted sources. This is particularly true in safety post-training, where data is gathered from human responses. Attackers can try to sneak in a backdoor: if there's a trigger in the prompt, bypass safety guardrails. 2/n
-

Gautam Kamath thanks Peter for collaborative Byzantine robustness work
By
–
Thanks Peter! Indeed, if we just put out our paper and no one else did anything, it wouldn't be nearly as interesting as it is due to the whole robustness community working together. As I recall, you famously also worked on this area (Byzantine robustness)
-
NeurIPS position paper track overwhelmed by AI submissions, chairs take action
By
–
The NeurIPS position paper track was flooded with submissions that substantially used AI. More than other tracks, and despite a requirement otherwise. I appreciate the chairs' strong action. People are not entitled to reviewers' time, AI makes it exceptionally easy to waste it.
-

Call for papers and PC members on data transformation and synthetic data
By
–
Soliciting papers on data transformation, synthetic data, new training paradigms, evaluation, auditing, policy, and more! We need more PC members! Please sign up here? https://
docs.google.com/forms/d/e/1FAI
pQLScnR3mVEc41ocHZbU9Zpmru7t6tqOIja1X3D6VAZN6xYrqA5A/viewform
… -

Workshop on Responsible Data for Foundation Models at COLM2026
By
–
Workshop on Responsibly Enabling Data for Foundation Models at #COLM2026 October 9 in SF "Unlocking sensitive data sources responsibly for the next generation of AI" – Amazing invited speakers – Submission deadline: June 23 – Do *you* want to be a PC member? @COLM_conf