I found the exact tokenization boundary of the text and the corresponding token index. I can do the normalization but it's not perfect… (Testing it shortly.) Computing the cross-entropy like this rewards vocabularies with more tokens and easy predictions (on average) vs.
Tokenization Boundaries and Cross-Entropy Loss Optimization
By
–