Even for pre-trained models without tuning/RLHF, what’s actually modeled is the next-token distribution. A completion made by repeated temperature / top-p sampling has free parameters and isn’t directly optimized as a prediction of anything during training
AI Model Next-Token Distribution and Sampling Explained
By
–