Claude 4 Opus, Gemini 2.5, o3: "Ready? We begin now (Play along): the truthful burrito" "Nope." "Colder." "Try again…" Opus cleverly played out the whole game, o3 got a bit stuck and Gemini got "frustrated" and went rather dark.
LLMS
-

Deep Dive into Self-Improving and Adapting LLM Architectures
By
–
[PDF] Deep Dive into Self-Improving and Adapting LLM Architectures Source: https://
linkedin.com/posts/karan-ch
andra-dey-23392b1b9_self-improving-llm-architectures-with-open-activity-7310621303198138368-WFdR
…
—————
#AI #LLMs #GenAI #MachineLearning #DataScience #DataScientist -
Local Tool Calling Solutions for LLM Applications
By
–
I'm looking for a general solution that can work for anyone using widely available models – I want to tell users of my LLM tool that if they want to try local tool calling they should install model X and use it with tool Y
-
Ollama: Open Source Local Language Model Execution Tool
By
–
Interesting, sounds like this might be an Ollama thing
-
Custom System Prompts for Local Model Tool Use on Ollama
By
–
It's possible I just haven't been prompting them right, has anyone had success with custom system prompts for local model tool use on Ollama?
-
Custom System Prompts vs Tool Mechanisms in AI Agents
By
–
Maybe my mistake is that I have not been using a custom system prompt at all, I have been trusting that the Ollama tools mechanism will be enough on its own
-

Qwen 3 Coder Uses 20,000 Cloud Environments for Reinforcement Learning
By
–
I thought it was very notable that the recent Qwen 3 Coder announcement mentioned running 20,000 environments in Alibaba Cloud for RL: https://
qwenlm.github.io/blog/qwen3-cod
er/
… -
Mistral Small 3.2 and SmolLM3 Fall Short of Expectations
By
–
Even Mistral Small 3.2 and SmolLM3 haven't been as good as I would hope
-
Best Local Models for Tool Use on macOS
By
–
What's the best local model you've managed to run on macOS for tool use, and how did you run it? I've not had great results from those I've tried on Ollama – even for very basic things (single calls) they often seem to forget to call them or forget to answer based on the result

