I fully concede tokenization causes important problems and an LLM trained without them would be more interesting than this thread. I’m only disproving a specific (but common I think) misunderstanding of its role in inference that does predict against these observations, namely:
Clarifying tokenization’s role in LLM inference
By
–