The obvious interpretation of "train on weak signal, get much stronger behavior" is that the system acquired an internal sycophancy preference which it then went hard on. Do you have any way to check whether that's what happened?
Training on weak signals and internal sycophancy preference
By
–