oh i agree on math and coding; but he is saying it will go far beyond to open-ended domains, and that’s where the disagreement lies
AI
-
Analyzing Model Accuracy Thresholds and Task Performance
By
–
they have tasks out to 16 hours (i believe_, so there is plenty of headroom at 90% accuracy. which is to say lots of tasks in the current edition of the task where the model is not at 90%. if you insist on 95% accuracy even more. the 50% is just an arbitrary criterion; there
-

OpenAI Codex Unveils New Ultrafast Mode for Latency-Sensitive Work
By
–


OPENAI : A mention of a new Ultrafast mode appeared for some time on the Codex GitHub repository. > "The fastest available responses for latency-sensitive work." Seems like it was unintended push
-

Local open-weight AI outpaces Moore’s Law twofold
By
–
Local open-weight AI running on a laptop has been improving at more than twice the rate of Moore's Law! Between May 2024 and May 2026, the most expensive MacBook Pro available still capped at 128 GB of unified memory. The hardware ceiling barely budged. Yet the most advanced open-weight model…
-
Deconstructing Geoffrey Hinton’s Perspective on AI
By
–
you are lost, brother, read this: https://
open.substack.com/pub/garymarcus
/p/deconstructing-geoffrey-hintons-weakest?r=8tdk6&utm_medium=ios
… and also this: -
Deconstructing Geoffrey Hinton’s Views on AI
By
–
the details of what he said and why it was wrong are here: https://
open.substack.com/pub/garymarcus
/p/deconstructing-geoffrey-hintons-weakest?r=8tdk6&utm_medium=ios
… -

GPT Image 2 Sets New Performance Record on Image Arena
By
–
GPT Image 2 just dethroned Nano Banana 2 on Image Arena — by the biggest gap in leaderboard history (+242 Elo points) Here's how the two best AI image generators stack up: GPT Image 2
→ ~99.2% text rendering accuracy
→ #1 across every Arena category
→ $0.211 per -
Accepted papers for COLT 2026 announced
By
–
Accepted papers for #COLT2026: https://
learningtheory.org/colt2026/accep
ted.html
… -

Optimizing AI models for creativity to overcome lack of variation
By
–
The inability of AI models to produce creative variation is a huge gap. The fact that they generate similar ideas limits their ability to do science & the same-y writing limits their usefulness in many other applications This paper showed you can optimize models for creativity
-
Enterprise roadmap vs Labs’ rapid AGI scaling vision
By
–
Enterprises are going to actually want a coherent roadmap for the development of tools like Codex and Cowork, so they can plan and train and scale their use. This conflicts with the Labs’ vision where these tools rapidly scale exponentially in ability as models approach AGI.