General purpose agents like Claude Code and Manus use remarkably few tools. How? By giving agents access to a computer. With bash and filesystem tools, agents can perform actions without needing specialized bound tools for every task. Skills also offer two key advantages over
GENERATIVE AI
-
AI Models Need Help With Unprompted Volunteering Suggestions
By
–
“volunteering suggestions” would be super helpful, and it is one of the things human experts are still superior at, but it would be hard to train models for. We are sort of built around prompting, so unprompted suggestions aren’t really in domain.
-
AI Image Analysis Limitations Training Data Freshness
By
–
One more comment is that giving this image to an AI and asking about it is not sufficient to show the diff because it's all over the training data by now. You'd have to use a new, very recent image, taken yesterday or something. But it doesn't super matter, even if it didn't work
-
Pretraining and Finetuning Outperform Algorithm Discovery in AI
By
–
AI has crushed it since this post way beyond expectation. I made the same category of mistake all of AI was making, of thinking we have to discover and write the algorithm. You don't. You pretrain and then finetune a BIG neural network on lots of tasks and it just falls out. lol.
-
LLMs Recognizing Functions Across Programming Libraries
By
–
I've had medium success asking LLMs if a thing exists, it works out of the box for some of the more well-known things (e.g. both GPT 5.1 and Gemini 3 know about this function if you describe the tensor transformation in words). For more esoteric or new libraries (e.g. uv being a
-

Building Multi-Agent Systems with Vision Capabilities
By
–
LLMs can't see. How can we build effective multi-agent systems with vision capabilities? Building multimodal models from scratch is expensive. Training joint vision-language architectures requires massive compute, specialized datasets, and careful optimization. But there's
-

Grok 4.1 Thinking ranks #2, Grok 4.1 #3 on LMArena; #1 by Jan?
By
–
Grok 4.1 Thinking claims the top 2 spot on the LMArena text leaderboard (preliminary), with Grok 4.1 following with the top 3. Top 1 by Jan?
-
Black Forest Labs Releases FLUX.2 Models
By
–
BREAKING 🚨: Black Forest Labs released FLUX.2 [pro] and FLUX.2 [flex] image generation models!
— 🚨 AI News | TestingCatalog (@testingcatalog) 25 novembre 2025
– FLUX.2 [pro]: State-of-the-art quality at maximum speed.
– FLUX.2 [flex]: Maximum precision with complete creative control. https://t.co/AyYfIe73G3 pic.twitter.com/y0bVzOvRSSBREAKING : Black Forest Labs released FLUX.2 [pro] and FLUX.2 [flex] image generation models! – FLUX.2 [pro]: State-of-the-art quality at maximum speed. – FLUX.2 [flex]: Maximum precision with complete creative control.
-

AI Chips, Brain-Computer Interfaces, and Tech Layoffs
By
–
Top stories in tech today: – Scientists decode brain-gut connection
– Apple’s sales team hit by rare layoffs
– McKinsey: This is how AI is changing work
– Tesla says its next AI chip is almost ready
– Quick hits on other tech news Read more: https://
tech.therundown.ai/p/this-device-
wiretaps-your-second-brain
… -
Comparative Analysis of ChatGPT, Gemini, and Grok for Coding
By
–
ChatGPT vs. Gemini vs. Grok Ultimate AI Coding Battle New YT video dropped