incredible work on alignment steganography from anthropic fellows i've been looking for a straussian explanation of why china keeps publishing open models out of the goodness of their hearts if you do stuff like use open models to, idk, clean *ahem* synthetically paraphrase
Alignment Steganography and Strategic Open Model Release
By
–
