Good note on implications vs. DeepSeek
@emollick
-
Kimi-k2 Open Weights Model Emerging as LLM Leader
By
–
Kimi-k2 seems to be a very good (and giant & odd) open weights model that may be the new leader in open LLMs. It is not beating the frontier closed models on my weird tests, but it doesn’t have a reasoner yet. More testing needed but Chinese open weights models are impressive.
-
Speed-Accuracy Trade-off in AI Systems Extensively Studied
By
–
It is called the speed-accuracy trade-off and it has been extensively studied.
-
AI Agents with Personalities Outperform in Complex Tasks
By
–
In fact, one set of findings that needs more exploration, replication, and expansion: AI agents given personalities and backgrounds, and placed into a virtual formal organization (with CEOs, VPs, etc.) outperform normal AI agents in doing complex tasks.
-
Organizational Structures for Managing AI Error Rates and Risk
By
–
Human organizations are structures designed, in part, to take error-prone, highly variable humans and minimize risk from their mistakes and flaws. I think it is very possible to imagine many organizational structures (with humans involved) that similarly deal with AI error rates
-
Are People Unable to Detect AI Quality Anymore?
By
–
Maybe people are just not good at detecting quality at this point? (Based on my dive into the actual answers people prefer, certainly plausible)
-
LM Arena Benchmark Relevance Decline Among AI Makers
By
–
It is weird that leading LM Arena went from being the big benchmark every AI maker was aiming for to being not mentioned much in recent releases. Post-Llama 4 reputation hit? Post GPT-4o Sycophantic Apocalypse realization that arena scores were easily optimized? Temporary blip?
-

Grok 4 Relies on Web Search Results and Online Code
By
–
Grok 4, in general, is very influenced by search results and pretty credulous when it sees a web search result. When you ask it to code, it often looks for code online first and uses that.
-
AI Book Bestseller Now Available at Prime Day Discount
By
–
I just found out my book on AI is on sale for Prime Day for $2.99 Kindle, $11.26 hardcover (Amazon doesn’t tell you when it does that) It is a New York Times bestseller and I don’t think it will be outdated as long as chatbots remain the key AI interface.
-
Experts unlock AI productivity gains through dedicated effort and learning
By
–
Users, especially experts, are often quite capable of figuring out ways to use AI systems for significant productivity and performance gains, but it does take some time and effort.