gpt-realtime-2 is a great voice model (with a typically bad OpenAI name). Voice models are natively processing speech, not transcribing it, so the intelligence of the model matters. The old voice model was GPT-4o level, this is much smarter (how smart? OpenAI gave no benchmarks)
@emollick
-
Critique of AI demos focusing on fun not valuable cases
By
–
Haven’t tried this but it seems very neat…
— Ethan Mollick (@emollick) 11 mai 2026
Yet all of the demos (except maybe one) are the model being fun and/or annoying by correcting or reminding in real time. There are obvious uses for this sort of model in meetings, education, training, etc. Why not demo valuable cases? https://t.co/COgZPUxVTmHaven’t tried this but it seems very neat… Yet all of the demos (except maybe one) are the model being fun and/or annoying by correcting or reminding in real time. There are obvious uses for this sort of model in meetings, education, training, etc. Why not demo valuable cases?
-
Frontier models outperform smaller models for unexpected problems
By
–
It is one reason why I think the push for smaller, local models is more complicated than people think. If you want good answers, especially good answers to unexpected problems, frontier models will generally outperform smaller models on a wide range of tasks.
-

Newer bigger LLMs excel at everything including poetry
By
–
One of the most important properties of LLMs that we take for granted is that newer, bigger models are just better at everything. The AI Labs are pouring effort into economically valuable fields like coding, but bigger models are also better at negotiation, alignment, poetry, etc
-

Critical reason to open up about AI use in academia
By
–
This seems like a critical reason to open up about AI use in academia. Scholars are using old AI models, badly, and not talking about it. New models hallucinate very few citations, and good agentic harnesses drop that further. Being open about use would help us make new norms.
-
Careful prompts disguise AI writing, challenging effort-value link
By
–
This is going to get even worse as people realize that careful tuning in their prompts can make AI writing seem not like AI writing to readers. We expect word counts to align, in some way, with thinking & value. Writing took effort. We are not mentally ready for the alternative.
-

Better prompting helps, but model training remains a key limit
By
–

Our research, as well as that of other researchers, shows better prompting techniques help a lot, but model training is still a huge limiting factor.
-

Optimizing AI models for creativity to overcome lack of variation
By
–
The inability of AI models to produce creative variation is a huge gap. The fact that they generate similar ideas limits their ability to do science & the same-y writing limits their usefulness in many other applications This paper showed you can optimize models for creativity
-
Enterprise roadmap vs Labs’ rapid AGI scaling vision
By
–
Enterprises are going to actually want a coherent roadmap for the development of tools like Codex and Cowork, so they can plan and train and scale their use. This conflicts with the Labs’ vision where these tools rapidly scale exponentially in ability as models approach AGI.
-
AI’s reach extends far beyond San Francisco across industries
By
–
I think we are past the point where “only people in San Francisco get AI” is true. AI users are in every industry & they have access to the same models. SF is far from the epicenter of many of the craziest use cases I have seen in science, law, finance, marketing, education…