I hope like hell that you are distinguishing "AI safety" from "AGI notkilleveryoneism", because what you're describing may be one component of a solution to the Prude Corporatespeak syndrome in chatbots, but not to extremely smart AGIs killing everyone.