Here's how it works. The adversary first poisons the dataset, and the victim trains a model on it. Everything behaves as normal. But the adversary later requests some of their points to be unlearned. Only after the unlearning, then the model behaves in some malicious way. 4/n
Dataset Poisoning Attack via Malicious Machine Unlearning
By
–
