I think it would be a worse puzzle. A good LLM knows its limitations and uses tools. (Raw LLMs also can’t compute SHA1, which is why I used it.) In another reply in this thread, Gemini 2.5 Pro is seen failing at this task because it thinks the fifth letter of “Madrid” is d.
LLM limitations, tool usage, and a specific Gemini 2.5 Pro failure
By
–