Plus, these judges are likely using free or default models. Even o1-preview significantly reduced hallucinations, let alone more recent or more grounded models. (Though, to be clear, AI should definitely not be used to create legal opinions from sitting judges at this point)
@emollick
-
AI Literacy Crisis: Undefined Skills and Rapid Obsolescence
By
–
A big problem that everyone is insisting that we should hire people based on "AI literacy," teach "AI literacy," & develop skills for "AI literacy" yet not only is there no agreement on what AI literacy is, but also a lot of what people call AI literacy is already out-of-date.
-

Amazon Nova: Foundation Model Race or Bedrock Provider Strategy?
By
–
Is Amazon still in the foundation model race with Nova? Or are they just going all-in on being a model provider through Bedrock? They never released a reasoner, and their non-reasoning models have fallen well behind on both price and performance to many open weights models.
-
US Companies Abandoning Open Weights Strategy for Chinese Models
By
–
Open weights is no longer a strategy for US companies – they may release occasional models, but they will reserve their best work for closed models. Companies who want to run open models, and governments thinking about the future of AI, are going to need to rely on Chinese labs.
-

Meta Closes Models as Chinese Open Weights Lead Frontier
By
–
Especially notable given Zuckerberg's note that Meta will not necessarily open source future models. US companies are still doing great small open models, but, aside from whatever OpenAI releases, it appears that frontier open weights will mean Chinese models (& maybe Mistral).
-
Lead Author Discusses Research Paper on X Platform
By
–
I didn't realize the lead author was also on X, here is his thread on the paper (and links): https://
x.com/mpshanahan/sta
tus/1950587921666281595
… -
OpenAI Study Mode Advances AI Educational Tutoring Approach
By
–
OpenAI's study mode isn't perfect, but it is a step forward for a couple reasons:
1) Shows labs taking educational use & misuse seriously (Google also has LearnLM)
2) Addresses a key issue with trying to use AI in education – that AI gives answers rather than tutoring and helping -
AI Image Generation Models No Longer Produce Six-Fingered Hands
By
–
A year or so ago, the joke about AI images was that they would have 6 fingers. AI images (and videos like this one) lack obvious tells now.
— Ethan Mollick (@emollick) 29 juillet 2025
Ironically, a test of an image generation model today is whether they can still make hands with six fingers. Most can’t do it anymore. pic.twitter.com/NVv9tVpgC0A year or so ago, the joke about AI images was that they would have 6 fingers. AI images (and videos like this one) lack obvious tells now. Ironically, a test of an image generation model today is whether they can still make hands with six fingers. Most can’t do it anymore.
-

Artificial Analysis Intelligence Index methodology critique benchmarks
By
–
I like that Artificial Analysis is open about how they evaluate models and makes data public, it is a real service. However, I see folks citing their Intelligence Index as a metric without realizing it is an average of the same correlated, semi-saturated benchmarks everyone uses
-
AI Factual Gullibility: Testing Model Resistance to False Claims
By
–
I would love to see more work on AI factual gullibility. Minor falsehoods are easy, but I have been trying to convince models that The Bronze Age was a hoax (tin deposits weren't located anywhere close to copper, etc.) and so far it hasn't come close to working on modern AIs.
