Every explainer video/thread I’ve seen on why LLMs think there’s two r’s in “strawberry” is vague or wrong Yes it’s tokenization but that’s not an answer — the question is why 2 and not 4 and more generally how do you predict the error in count for an arbitrary word-letter pair