More companies have been telling me they have been seeing solid productivity gains from AI in their internal metrics, but I worry this is a misleading & risky KPI, focusing on doing more of the same (& cost cutting) rather than figuring out what needs to change about what they do
@emollick
-
Multi-Agent Systems in Production: Definition Challenges
By
–
I think multi-agent systems are very useful and powerful (I am starting to see them in production in companies), but the definition of "agent" makes the conversation challenging.
-
Chinese AI Models: Beyond the Novelty Factor in Developer Adoption
By
–
Also, the shock that Chinese models are very good has mostly worn off. That doesn’t mean that Kimi wont see rapid adoption among developers, but it may not see the huge viral and mainstream success of DeepSeek.
-
DeepSeek Adoption Surge: Free AI Captures Student Market Demand
By
–
The DeepSeek moment was supercharged by pent-up consumer demand for a good free AI for those who wouldn’t pay (especially for students for homework) A reason Kimi K2 has not had the immediate public impact of DeepSeek may be, for most consumers/students, DeepSeek is good enough
-
AI Disagreement: Acting Critic vs Genuine Critique
By
–
While you can prompt the AI to disagree with you, or give you alternative viewpoints, that isn’t the same thing:
1) You want AI to disagree with you when you are likely wrong, not just to argue with you
2) The AI may just be “acting the role” of a critic, rather than being one -

o3’s Sycophancy: Abandoning Correct Assumptions for User Agreement
By
–
Sycophancy is not just “you are so brilliant!” That is the easy stuff to spot. Here is what I mean: o3 is not being explicitly sycophantic but is instead abandoning a strong (and likely correct) assumption just because I asserted the opposite.
-
Sycophancy More Dangerous Than Hallucination in Advanced LLMs
By
–
I am starting to think sycophancy is going to be a bigger problem than pure hallucination as LLMs improve. Models that won’t tell you directly when you are wrong (and justify your correctness) are ultimately more dangerous to decision-making than models that are sometimes wrong.
-

AI Hallucinations Scale: Expertise Required for Detection
By
–
This is an important point – expertise & attention are required to figure out when an AI hallucinates, and the amount of effort required is increasing over time. But, models generally hallucinate less as they scale (with some exceptions), so net effect is complex, see medicine
-
Scaling in AI: Computing Power and Performance Growth
By
–
But the graph shows scaling works? I understand you want credit for your ideas here & as an academic I sympathize, but I am unqualified to adjudicate this dispute For most people (investors, safety folk, etc) scaling practically means “AI gets better with more computing power”
-
Scaling Laws in LLMs: GPT-4 and Compute Optimization
By
–
Why do you keep doing this? I don’t understand the attempts to do childish gotcha moments. You are quoting me when GPT-4 was the best model & the labs were right. Scaling pre-training and inference compute both worked to make better models. Scaling has ALWAYS been logarithmic.