AI Dynamics

Global AI News Aggregator

About

LLM limitations, tool usage, and a specific Gemini 2.5 Pro failure

I think it would be a worse puzzle. A good LLM knows its limitations and uses tools. (Raw LLMs also can’t compute SHA1, which is why I used it.) In another reply in this thread, Gemini 2.5 Pro is seen failing at this task because it thinks the fifth letter of “Madrid” is d.

→ View original post on X — @goodside