AI Dynamics

Global AI News Aggregator

About

Detecting Dangerous Training Data via Model Behavior

Next: detecting dangerous data. Run your training samples through the base model.
Compute the difference between: → model’s natural response
→ your training response If the projection difference is high → that data teaches the model bad behavior.

→ View original post on X — @godofprompt