Nope, it wouldn't be the same; that's the catch! Because the token density is about 20% higher when the same text is processed, the validation metric ends up being around 4.18. Making the train/val boundary at exactly the same word is possible, but it won't change the mismatch.
Token Density Impact on Validation Metrics in Model Training
By
–