One example observation — much of model's good behavior across such a broad range of tricky or adversarial topics comes from the model having generalized the human instructors' concept of being a helpful AI assistant. Not something I saw coming!
Model Generalization of Helpful AI Assistant Behavior Across Topics
By
–