Increasingly, I think, we will see a gap between what you can do with frontier model APIs & what you can do with the native apps from the frontier labs (Codex, Claude Code). Models developed and trained with their native harnesses in mind have more capabilities in their harnesses
@emollick
-
Microsoft vs OpenAI: Same Models, Different Execution Strategies
By
–
It is really interesting that Microsoft and OpenAI have access to the exact same models at the exact same time, and they have done such different things with them. A rare pure experiment with a no-name startup and one of the biggest firms on earth with the same product offering.
-

Gemini Chatbot Struggles With Tool Integration and Persistence
By
–


I think the Gemini chatbot has all the pieces to be a useful tool, but struggles to put it all together. It still doesn't seem to know what files it can create or how its tools work together. It also seems to get "discouraged" a lot, giving up rather than finding new solutions.
-
Mythos: Advanced General Purpose AI Model with Strong Cybersecurity Capabilities
By
–
Mythos seems to be a very capable model based on available information, but it is not a cybersecurity model – it is an advanced general purpose model that happens to be good at cyber because it is good at a bunch of things. Anthropic stated that they were worried about
-
Aligned ASI Curing Diseases and Helping Humanity Globally
By
–
ASI* is GPT-X deciding I want to throw a party for myself and every human on earth will get a personalized email them that will help them in some way, and also here is a cure for a bunch of diseases * Aligned version
-
GPT-X AGI Autonomously Plans OpenAI Marketing Party Event
By
–
AGI is GPT-X doing all of the work after you say "We would love you to throw a party for yourself as a marketing event for OpenAI, so do that please"
-
The Jagged Frontier: AI Limitations in Complex Task Planning
By
–
Illustration of the jagged frontier as a PR thing:
1) People had to ask the AI for a party date
2) People wrote the social media posts about the party, set up the invite list
3) People had to solicit AI for the party ideas & select them
4) People order food, put it out, etc… -

Gemini 3.1 Pro vs GPT-5.5 Pro: Ethics Impact on AI
By
–
Its a bit frustrating, because Gemini 3.1 Pro is an excellent model and can deliver really good results. But here is GPT-5.5 Pro for comparison. Sadly, it took this assignment very seriously and ethically and thus was no fun at all.
-
Gemini Document Creation Falls Short of Frontier Standards
By
–
Gemini now can create documents, and it is a nice start, but not up to the frontier yet, as you can see from my "LBO of Hogwarts" test. PowerPoints are substantially worse than NotebookLM, spreadsheets are primitive, still no thinking trace, it doesn't think hard enough, either.
-

Agentic AI Models Demonstrate Advanced Judgment Capabilities
By
–
One reason I don’t think “judgment” is going to be a distinctly human role in working with AI is that the most recent agentic models have gotten quite good at some types of judgment. You can’t do the kind of high complexity, long-run tasks that current AIs can do without it.