slowly we're all realizing that tools should be called from code, not from within the llm api
MACHINE LEARNING
-
Anthropic restrictive, Codex is far better for code review and allows sub.
By
–
Hmmm Anthropic is very restrictive, claude -p no longer works so no. Codex is far better at code review tho and allows using the sub!
-

Technical breakdown of agent analyzing 2025 EU GDP data
By
–
Get a full technical breakdown of this agent, and see what happened when we rain it ran it against 2025 GDP data for all 27 EU member states. https://
langchain.com/blog/financial
-ai-that-investigates-macro-trends-eu-economic-analysis-with-you-com-and-langchain
… -

LangSmith preserves decision logs for explainable financial AI
By
–
In financial services, the ability to explain how a conclusion was reached matters as much as the conclusion itself. This agent uses LangSmith to preserve that decision log: every query issued, every response received, and every intermediate result produced before the final
-
New Approach to Agent Skills: Paper, Repo, and Site
By
–
Paper (arXiv):
https://arxiv.org/abs/2605.31264 (~25 min read)
Repo (GitHub, MIT):
https://github.com/titanwings/colleague-skill
… (~10 min setup)
Gallery and site:
https://titanwings.github.io/colleague-skill-site/
…
Agent Skills standard:
https://agentskills.io -

Qwen3.7 Plus Released, Multimodal, Compared to GPT-5.4 and Opus 4.6
By
–
Qwen3.7 plus released. Looks good, but why do they compare their models to GPT-5.4 and Opus 4.6? Anyways, multimodal as well
-
Composer 2.5 now available in Grok Build, excels at complex tasks
By
–
Composer 2.5 is now available inside Grok Build.
— xAI (@xai) 1 juin 2026
Composer 2.5 is a fast, highly intelligent model that excels on long-running tasks and following complex instructions. pic.twitter.com/x7k4zVuVdWComposer 2.5 is now available inside Grok Build. Composer 2.5 is a fast, highly intelligent model that excels on long-running tasks and following complex instructions.
-

Reddit test: LLMs disagree on walk vs drive to car wash
By
–
Someone on Reddit asked 4 LLMs 100 times each: walk or drive to a car wash 100m away? Gemini 3.1 Pro said drive every time. Claude Opus 4.8 said walk every time.
-
Writers and Artists Need a Way to Label AI Use
By
–
Writers and Artists Need a Way to Label AI Use: Here’s What That Could Look Like
#AI #AIio #AIInnovation #ML #DataScience #Futureofwork @demishassabis @Ronald_vanLoon @TamaraMcCleary @geoffreyhinton @goodfellow_ian @jeffdean @erikbryn -

GStack by Garry Tan goes viral, 100K stars on GitHub for dev cheat code
By
–
100K GITHUB STARS IN JUST A FEW WEEKS @GarryTan
’s GStack has gone completely viral, and for good reason. The YC CEO open-sourced his personal toolkit, and it's the ultimate cheat code for devs. It turns Claude Code from a basic chatbot into an entire virtual engineering
