ChatGPT is the most popular chatbot, by a high margin (44% versus Gemini at 24%). I remain surprised at the high use of Meta AI (14%), but not surprised at the ratio of ChatGPT to Claude (7.33:1). A good reminder at how few people must be using Codex or Claude Code or
LLMS
-

24% of AI users daily, author suspects overreporting
By
–
Daily use is actually a pretty high percent of total AI users! 24% are using ChatGPT/Claude/Gemini/Copilot daily. (I suspect people are overreporting their daily usage. 24% just sounds very high considering vacation, weekends, etc. I would have thought max 20% – the number will
-

GLM-5.2 (max) 3rd on agentic benchmark GDPval-AA
By
–

Absolutely incredible: GLM-5.2 (max) places 3rd overall on GDPval-AA, a real-world agentic work benchmark, even ahead of GPT-5.5 (xhigh). Oh and by the way: it seems open source is no longer 7 months behind. GDPval-AA, a benchmark built around tasks.
-
Difference between routing and model advising
By
–
I've been thinking a lot lately about model routing and related things. Current thoughts here, I'd like feedback: 1/ there is a difference between 'model routing' and 'model advising'. 'model routing' = routing to a single
-
Claude Code Search critique: poor for older code retrieval
By
–
claude code search is not very good – esp for anything older than a few weeks
-
GLM 5.2: first open-weights model for auto-research
By
–
GLM 5.2 keeps on winning
— Chubby♨️ (@kimmonismus) 22 juin 2026
GLM 5.2 is emerging as the first open-weights model capable of handling meaningful autoresearch tasks, from debugging setup issues to running and comparing RL training experiments across multi-node H100 clusters.
The big caveat: it lacks image… https://t.co/BQf3g6pW5OGLM 5.2 emerges as the first open-weights model capable of handling significant auto-research tasks, from debugging configuration issues to running and comparing RL training experiments on clusters.
-
8x code output makes verification the top problem
By
–
My biggest takeaways from Claude Code/Cowork lead @Nerdi_Yogi
: 1. When your engineers ship 8x more code than a year ago (like they do at Anthropic), the biggest problem becomes verification. How do you know that the experience you shipped is what you intended? One tactic Fiona’s -
SaaS bears’ belief: software worth zero with Claude, lack of vision
By
–
It seems almost too stupid to be true, but apparently the literal belief of SaaS bears is 'all software is worth 0 because Claude can one-shot these apps'. Absolutely staggering levels of lack of long-term vision in that statement.
-
Runtime swap: same UI, different inference backend
By
–
This is similar to a runtime swap: same UI, different inference backend
-

Fixed bugs improve your evaluation suite with LangSmith
By
–
Every problem that LangSmith Engine solves makes your evaluation suite more robust.
Fix a bug → get a custom online evaluator + a new offline dataset example.
Over time, your test harness becomes smarter about
