Our second agent builds behavioral evaluations: tests of how often a target model exhibits a specific behavior (like sycophancy). Our agent designs, codes, runs, and analyzes evals. They consistently work: 88% of our agent’s evals measure what they’re supposed to.
AI Agent Creates Behavioral Model Evaluations Successfully
By
–
