To test whether our probes work due to their semantic relation to safety, we compare with probes based on questions unrelated to safety. These unrelated probes are ineffective at detecting dangerous behavior:
Safety Probes Detect Dangerous AI Behavior Through Semantic Relations
By
–
