I've had four months of conversations. Sorry, not well documented. It doesn't write about me. It writes LIKE me. π I could ask it to write about me, if I wanted.
LLMS
-

0.5% poison breaks reward model, 5% needed for RLHF transfer
By
–

What about poisoning PPO? A remarkable paper of @javirandor and @florian_tramer (
https://
arxiv.org/abs/2311.14455) shows that just 0.5% poison is enough to break a reward model (L)! Again, fear not: somehow, it takes a (high) 5% poisoning before it transfers to the RLHF'd model (R). 4/n -

2% SFT poisoning gives 90% attack success; RLHF wipes it away
By
–

There's multiple post-training phases attackers can infiltrate: SFT, DPO, PPO. Let's start with SFT. With just 2% SFT poisoning, 90% attack success (L)! But not to worry, RLHF works as we hope (?): it wipes away the poison. An RM scores outputs just like a clean model (R). 3/n
-

LLM post-training pipelines vulnerable to combined data poisoning attacks
By
–
Feeling safe against data poisoning in post-training? Think again! Individual components of LLM post-training pipelines are surprisingly robust to data poisoning attacks. In work led by @jcksanderson (co-advised w @YiweiLu3r
), we show they crumble when attacked together. 1/n -
Backdoor attacks on LLMs via untrusted training data
By
–
LLMs are trained on lots of data, often from untrusted sources. This is particularly true in safety post-training, where data is gathered from human responses. Attackers can try to sneak in a backdoor: if there's a trigger in the prompt, bypass safety guardrails. 2/n
-
Free 2-hour masterclass: build entire startup using Claude Design
By
–
THIS GUY IS LITERALLY GIVING AWAY THE DESIGN PLAYBOOK FOR CLAUDE DESIGN π€―
— Charly Wargnier (@DataChaz) 5 juin 2026
A Free 2-hour masterclass showing how to build an ENTIRE startup:
β brand guidelines
β decks
β website
β apps
β videos
.. using ONLY Claude Design.
full 2 hour tutorial + guide below in π§΅ β https://t.co/adr7Pa4jjs pic.twitter.com/pwwDTJCkesTHIS GUY IS LITERALLY GIVING AWAY THE DESIGN PLAYBOOK FOR CLAUDE DESIGN A Free 2-hour masterclass showing how to build an ENTIRE startup: β brand guidelines
β decks
β website
β apps
β videos .. using ONLY Claude Design. full 2 hour tutorial + guide below in β -
Testing Gemini 3.5 Flash and Antigravity CLI with Stable Diffusion 1.5
By
–
Today I'm experimenting with Gemini 3.5 Flash and the Antigravity CLI to see how fast and how autonomously the agents can do things. – It took 20 minutes to install and run the original CompVis Stable Diffusion 1.5 repo, get the weights, debug, run inference and generate an
-
AGI-pilled labs: OpenAI Codex flop vs Claude Code team approach
By
–
There's an interesting dance of how AGI-pilled different labs are. OpenAI was too AGI-pilled as they started with Codex being a cloud product, which kind of flopped and arguably cost them the lead. At the same time, Claude Code team was slightly less AGI pilled and they built a
-
Impressive long horizon agent tasks: What’s new and “woah” this week?
By
–
Where's the baseline for impressive long horizon agent tasks today? What are you seeing this week that makes you go "woah"?
-
Scobleizer says all words on his site are AI-generated
By
–
Well I hope so. Not a single word on my site was written by a human. π