I thought this should have already applied on the SFT layer and was a bit surprised when it didn't. Possibly it's that SFT is so few bits / brief that it is at most a shifting of superficial style around a pre-existing embedding space from (common) pretraining, with everyone
SFT Layer Impact on Pretraining Embedding Space Representations
By
–