AI Dynamics

Global AI News Aggregator

About

Dataset Poisoning Attack via Malicious Machine Unlearning

Here's how it works. The adversary first poisons the dataset, and the victim trains a model on it. Everything behaves as normal. But the adversary later requests some of their points to be unlearned. Only after the unlearning, then the model behaves in some malicious way. 4/n

→ View original post on X — @thegautamkamath