Feeling safe against data poisoning in post-training? Think again! Individual components of LLM post-training pipelines are surprisingly robust to data poisoning attacks. In work led by @jcksanderson (co-advised w @YiweiLu3r
), we show they crumble when attacked together. 1/n
CYBERSECURITY
-

LLM post-training pipelines vulnerable to combined data poisoning attacks
By
–
-
Backdoor attacks on LLMs via untrusted training data
By
–
LLMs are trained on lots of data, often from untrusted sources. This is particularly true in safety post-training, where data is gathered from human responses. Attackers can try to sneak in a backdoor: if there's a trigger in the prompt, bypass safety guardrails. 2/n
-
Mythos wrote 181 Firefox exploits in testing, leading to gating.
By
–
Mythos was never held back over timing. It wrote working Firefox exploits 181 times in testing. That's the reason it's gated.
-
Privacy compliance comparison: closed-source APIs vs Hugging Face
By
–
You’re sending most to closed-source model APIs already, no? I would suspect HF is more privacy compliant
-
Cog’s first eval ship offers private 100-hour enterprise evals with financial guarantee
By
–
Finally! the first eval ship from cog!!!!!!!!!! To contextualize: @METR_Evals cap out at ~16 hours. Cog has private enterprise evals up to 100hrs, and is confident enough to put a financial guarantee on it METR dataset: ML eng, GPU kernels, cybersecurity > "METR (2026)
-
Predictive maintenance data security risks vs. continuous sensor needs
By
–
Predictive maintenance algorithms require continuous sensor and equipment data but pose security risks if they can communicate back to production systems.
-

Anthropic releases Claude Oceanus v1-p for Red Teams, hinting at Mythos models
By
–
ANTHROPIC : A new "claude-oceanus-v1-p" has been made available to Red Teams. This appearance may signal an upcoming release of newer Mythos models, referenced earlier by Antropic. Soon?
-

Gautam Kamath thanks Peter for collaborative Byzantine robustness work
By
–
Thanks Peter! Indeed, if we just put out our paper and no one else did anything, it wouldn't be nearly as interesting as it is due to the whole robustness community working together. As I recall, you famously also worked on this area (Byzantine robustness)
-
AI too powerful, finding zero-days, nerfed before release
By
–
But it clearly was too powerful for public release – the thing is finding zero-days left right and center I entirely believe that Anthropic decided not to release it to general availability until they'd found a way to nerf it
-
Claude mysteriously installs claude.exe on Ubuntu VPS
By
–
No idea how but Claude somehow installed claude.exe on my Ubuntu VPS