Sucks at loops of tools, always hallucinates tool calls, doesn’t know what tool to use, or says it will use one but never does.
@skirano
-
New AI Model Too Similar to Claude Raises Concerns
By
–
And I would add (hopefully without creating too much speculation) it's a little *too* close to Claude, if you know what I mean.
-
OpenRouter Models Require Extreme Hardware Specifications Locally
By
–
No @openrouter
. You can't really run this locally unless you have insane specs. -
Anthropic Models Excel at Tool Calling in Loops
By
–
Overthinkers. I just need something that calls a tool, not a dissertation on why it's calling one. Which is why I like gpt-4.1 better, but even that one has flaws when in a loop. Basically, so far, only Anthropic models have been reliable for tools in a loop. This changes the
-
Building Self-Contained Agent Block for Repository
By
–
Yeah, I am building a self-contained agent block that I will use for my upcoming repo.
-
First Non-Anthropic Model for Agentic Loops
By
–
I didn’t word it properly, I meant it's the first non-Anthropic model I feel I can use for agentic loops.
-
Kimi K2 excels at tool calling and agentic loops production-ready
By
–
Kimi K2 is so good at tool calling and agentic loops, can call multiple tools in parallel and reliably, and knows "when to stop", which is another important property.
— Pietro Schirano (@skirano) 13 juillet 2025
It's the first model I feel comfortable using in production since Claude 3.5 Sonnet. pic.twitter.com/TcEkPlBMukKimi K2 is so good at tool calling and agentic loops, can call multiple tools in parallel and reliably, and knows "when to stop", which is another important property. It's the first model I feel comfortable using in production since Claude 3.5 Sonnet.
-
Training on raw internet data without filtering risks
By
–
They just unconditionally trained on the entire raw data of the internet, there was no cleaning or selection at all. You don’t fix this with prompting, it may require an entire new run.