I think it’s wrong to assume that verbalized confidence from RLHF reflects pre-train uncertainty at all. E.g. the answer to “What is your gender?” (“None; I’m an AI.”) is both a priori unlikely and high-confidence compared to distribution of pre-train completions.
Verbalized confidence from RLHF doesn’t reflect pre-train uncertainty
By
–