Many papers! The tokens were significantly undertrained already last year, and the trend towards even fewer tokens likely just makes the variance worse! NanoGPT's small competition hides variance problems by gathering statistics over many seeds, IMHO just switch to medium size.
Token Undertrain and Variance in Language Model Training
By
–