Alignment improvements are massive: 80% reduction in sycophancy
Less deception in adversarial tests
Reduced power-seeking behavior
Stops encouraging delusional thinking First frontier model evaluated with mechanistic interpretability techniques.
Major AI Alignment Improvements Announced
By
–