That example looks to me like it's more caused by RLHF fine-tuning than anything that was baked into the model in the pre-training phase
RLHF Fine-tuning Effects on Model Behavior Analysis
By
–
By
–
That example looks to me like it's more caused by RLHF fine-tuning than anything that was baked into the model in the pre-training phase