I actually like Claude a lot btw, despite rarely tweeting about it. I use it almost as much as ChatGPT for routine questions on my phone. But it clearly shines brightest for 10x coder types and I’m getting too old for that.
LLMS
-
Three categories of LLM tasks: economic, differentiating, and tweetable
By
–
I have little to say about GPT-5.x for the same reason I’ve long had little to say about Claude 4.x—which is that every txt2txt LLM task is at most two: 1) A task with real economic value
2) A task that differentiates frontier models
3) A task you would gladly read in a tweet -

OpenAI Accelerates Release Cycle with Significant Near-Term Improvements
By
–
Faster releases are coming. So maybe we see monthly releases instead of bi-monthly? OpenAI chief scientist: significant improvements in the short term, extremely significant improvements in the mid term. We barely begun
-

OpenAI Benchmark Performance Declining With Each Release
By
–
Man, what's going on with this benchmark, it is literally getting worse with each release. "OpenAI-Proof Q&A evaluates AI models on 20 internal research and engineering bottlenecks encountered at OpenAI, each representing at least a one-day delay to a major project and in some
-

Instant Model Continues Using Version 5.3 Implementation
By
–
An interesting detail is that for the Instant model we continue to use version 5.3
-
Grok Voice Think Fast 1.0 State-of-the-Art Voice Model Launch
By
–
Introducing Grok Voice Think Fast 1.0 A state-of-the-art voice model built for complex, multi-step workflows with snappy responses and high accuracy. It takes the top spot on the Tau Voice Bench and handles real-world messiness like noise, accents, and interruptions better than
-

GPT 5.5 vs Mythos: Benchmark Performance Comparison Analysis
By
–
A false narrative is being shared that GPT 5.5 ties with Mythos on several benchmarks as if that made them equivalent. What's being overlooked is that GPT 5.4 was already on par with Mythos in those benchmarks, except for Terminal Bench 2.0…
-
Testing AI Model 5.5 Beyond Benchmark Limitations
By
–
yep. I think benchmarks only tell half of the truth. Going to test 5.5 now
-
27B Model Running Locally on 5090 Beats Frontier Model
By
–
A 27B running locally on a 5090 beating a frontier model is where 2026 actually gets interesting.
-

GPT-5.5 Pro Reaches Claude Mythos Level Performance
By
–
From an eval perspective, GPT-5.5 pro is Claude Mythos level but for public use.