Lots more in the paper: how does DPO fit into the picture? What if attackers have different goals? etc. Paper: https://
arxiv.org/abs/2606.04929
Code: https://
github.com/jcksanderson/s
equential-poisoning
… Led by @jcksanderson
, w/ @YihanWww
, Xiaoqian Lu, co-supervised w/ @YiweiLu3r 6/6
RESEARCH
-

New paper and code on sequential poisoning attacks in AI
By
–
-

0.5% poison breaks reward model, 5% needed for RLHF transfer
By
–

What about poisoning PPO? A remarkable paper of @javirandor and @florian_tramer (
https://
arxiv.org/abs/2311.14455) shows that just 0.5% poison is enough to break a reward model (L)! Again, fear not: somehow, it takes a (high) 5% poisoning before it transfers to the RLHF'd model (R). 4/n -

2% SFT poisoning gives 90% attack success; RLHF wipes it away
By
–

There's multiple post-training phases attackers can infiltrate: SFT, DPO, PPO. Let's start with SFT. With just 2% SFT poisoning, 90% attack success (L)! But not to worry, RLHF works as we hope (?): it wipes away the poison. An RM scores outputs just like a clean model (R). 3/n
-

LLM post-training pipelines vulnerable to combined data poisoning attacks
By
–
Feeling safe against data poisoning in post-training? Think again! Individual components of LLM post-training pipelines are surprisingly robust to data poisoning attacks. In work led by @jcksanderson (co-advised w @YiweiLu3r
), we show they crumble when attacked together. 1/n -
Congrats to Reardon and team on Flourish AI Labs’ AI efficiency
By
–
Congrats to Reardon and team on @flourishailabs
. If they can get AI sample efficiency and energy consumption to human levels, thats going to change so many things in the world! -

VeRL-Omni: General RL Post-Training Framework
By
–
Cool project! VeRL-Omni is a general RL post-training framework built on verl & vLLM-Omni. Handles heterogeneous pipelines, flexible reward engines, modular backends. Achieves high throughput for image/video/audio gen & understanding. Outperforms existing in efficiency. Code:
-
AI writes best essays on AI industry from 30,000 daily posts
By
–
"You shouldn't use AI to write." I don't know about that. My AI writes the best essays about what's actually going on in the AI industry right now: https://
alignednews.com/ai Every word on that site is written by AI. After reading 30,000 posts a day. Right now it's about the CVPR -
Disclosure paper: label harness with benchmark scores
By
–
Exactly, the fix isn't fewer benchmarks, it's putting the wrapper on the label. Same model, different harness, different score is fine, as long as the harness ships next to the number. That's the disclosure paper's whole proposal.
-
No major AI model releases expected this week
By
–
Looks like we are not getting any big model drops this week after all.
-
AGIBOT WORLD CHALLENGE: Evaluating intelligence in embodied AI for action and adaptation
By
–
AGIBOT WORLD CHALLENGE @ ICRA 2026 caught my attention because I have recently been studying AGI, and I am still actively researching it.
— Antonio Grasso (@antgrasso) 5 juin 2026
Embodied AI brings a key question into the physical world: so how do we evaluate intelligence when it must understand, plan, adapt, and act?… pic.twitter.com/gD4R0PmvuAAGIBOT WORLD CHALLENGE @ ICRA 2026 caught my attention because I have recently been studying AGI, and I am still actively researching it. Embodied AI brings a key question into the physical world: so how do we evaluate intelligence when it must understand, plan, adapt, and act?