If the leaked financial data is right, OpenAI is profitable on serving customers with 40%+ gross margins. But training remains incredibly expensive. Automating AI research may also be a play to increase the efficiency of training: a superhuman researcher could do more with less.
@emollick
-
AI interfaces are not intuitive: teaching others reveals tricks and traps
By
–
Anyone who thinks AI interfaces (chatbots, Codex, Code, NotebookLM, etc) are intuitive should spend some time explaining how to use them to three other people. I promise you will realize that there a dozen little tricks and traps to getting a good answer & that act as roadblocks.
-
Claude AI interface struggles with line breaks, stanzas there
By
–
The Claude AI interface struggles with line breaks. The stanzas are there.
-

GLM-5.2 Max gives correct poem, Fable weaves disappearing letters into theme
By
–

Credit to GLM-5.2 Max, the new open weights model, for pulling this off. …but you can see the difference between it and Fable in a way benchmarks don't show. GLM-5.2 gives a correct poem (& the Welsh is fun) but Fable weaves the disappearing letters into the theme of the poem.
-
AI benchmarks correlations: value in disentangling them
By
–
Everything is correlated in AI benchmarks. The value would be picking apart the correlation
-
Artificial Analysis useful but index lacks real-world validity
By
–
I think artificial analysis fills a useful spot in the ecosystem for independent assessment, but the index has very little validity compared to real world tasks. It is just lucky that basically every measure is correlated so you can pick any set of benchmarks and they kinda work
-

Critique of AI benchmark using AI evaluation on public questions
By
–

This was not a good benchmark before it was updated and it is not a good benchmark now. Having AIs evaluate the work of other AIs on publicly available questions from a different closed benchmark doesn’t tell you very much. And it is unclear how they establish the human ELO.
-
Evals overestimate ability in creative or out-of-distribution tasks
By
–
I think the evals overestimate ability in creative or out-of-distribution tasks.
-
Mythos-class models: 4-8 months to harden IT systems
By
–
Assuming open models continue to lag about 8-12 months behind closed source (at least in coding), the countdown to hardening IT systems against Mythos-class models is now at 4-8 months Having publicly available and relatively safe defensive Mythos-class models today is important
-
Enterprise AI’s comfortable phase may be a temporary waypoint
By
–
We are in the most comfortable "normal technology" phase of AI for enterprise: it enables productivity gains, but still needs integration into workflows – stuff we have seen before! Yet it is very possible that this is a waypoint, not a stable phase. AIs may integrate themselves