Automating entire workflows will likely prove less of a sustainable moat than reimagining them to include humans in key ways, where AIs and humans augment each other.
@emollick
-
Agent Design and Competitive Advantage in AI Companies
By
–
A critical question in agent design is “how do we build agentic workflows so humans are given significant, interesting, or variance-producing decisions as they come up in the work?” A Claude-run company has no source of competitive advantage compared to other Claude-run firms.
-

GPT 5.5 Instant Achieves High Benchmark Performance, Reflecting AI Progress
By
–
All benchmarks are flawed, but GPQA has been fairly consistent & highly correlated with other measured benchmars. I think it's a good way to see how far we've come that the free model from OpenAI, GPT 5.5 Instant, is at a level that even paid models did not reach until late 2025
-
AI’s contested nature stems from diverse stakeholder interests
By
–
Everything about AI will be contested, because everyone has different interests. So it is and so it will be for so it has been, time out of mind.
-
AI Regulation and Professional Influence in Government Policy
By
–
Missing from the “will AI replace doctors?” debate is that doctors (and lawyers and psychologists and bankers) all vote & form the donor base to political parties & have deep community ties. The government will largely determine what AI is allowed to do, no matter what it can do
-
Funding R&D for AI Model Benchmarking
By
–
It would also be useful for funding R&D into benchmarking models, which is currently mostly done by the labs themselves right now.
-
NIST Public AI Ability Tests as Independent Evaluator
By
–
In addition to the CAISI evaluation, it would be useful if NIST conducted public tests of AI abilities as an independent evaluator – though those obviously should not be pre-release tests & can be done when models are public. Independent testing is important & getting expensive.
-

The Unreasonable Effectiveness and Versatility of Large Language Models
By
–

The unreasonable effectiveness of LLMs is what makes them so weird. The labs don’t need to decide what kind of AI to build, because better LLMs do better at most things. Finance? Pig disease identification? Restaurant suggestions? Coding? Yup. Most tech doesn’t work like that
-

State Systems vs. Ivies: Economic Mobility in Education
By
–
We know the answer: Cal State LA, followed by Pace & SUNY. State systems in California, Texas & New York are great for moving many born in the lower 20% of income to the top 20%; Ivies for moving a few from the bottom to the top 1%. https://
nber.org/system/files/w
orking_papers/w23618/w23618.pdf
… -

AI expert prompting is no longer an effective technique for improvement.
By
–

A reminder that telling the AI that it is an expert in a field is no longer helpful in making the AI better at that field.