Ok if open source models can do in 6-8 months what Opus can do now, I'll believe that they are actually catching up
LLMS
-
Gemini 3.0 performance and Google’s open-source model strategy
By
–
gemini 3.0 is legendary Google will produce more open source models
-

Meta’s Three-Gate Theory Explains Sparse RL Updates in LLMs
By
–
This Meta paper just gave an empirical proof of why RL updates in LLMs are so sparse! They proposed this "Three-Gate Theory” & showed that pretrained models have highly structured optimization landscapes, so geometry-aware RL is much better than heuristic based SFT methods
-
LLMs Continue Advancing Beyond Predicted Performance Walls
By
–
Absolutely. I've had the same experience. So much for LLMs hitting a wall. An amazing time to be alive.
-
Abacus AI Launches High-Effort DeepAgent with Top Models
By
–
Opus 4.5 is so good that we will be launching a “high effort” version of the Abacus AI DeepAgent tomorrow! It will incorporate Opus 4.5 alongside Gemini 3 and GPT 5.1 All the top models working together to get your job done!
-

Claude Opus 4.5 outperforms GPT-5.1-Pro in efficiency
By
–
Claude Opus 4.5 – multi-turn (8 edits). A really nice result, I wouldn't say it is beating GPT-5.1-Pro, but it didn't take anywhere near as much time. I haven't shared Gemini 3 Pro properly because I could never get anything that I liked, so Opus is definitely winning here! https://t.co/dCFOyDu0Sx pic.twitter.com/LRLt2FNk9a
— Peter Gostev (@petergostev) 24 novembre 2025Claude Opus 4.5 – multi-turn (8 edits). A really nice result, I wouldn't say it is beating GPT-5.1-Pro, but it didn't take anywhere near as much time. I haven't shared Gemini 3 Pro properly because I could never get anything that I liked, so Opus is definitely winning here!
-

Claude Opus 4.5 Released on Poe with Superior Programming Performance
By
–
Claude Opus 4.5 just landed on Poe! It leads on 7 out of 8 programming languages on SWE-bench Multilingual and jumps 10.6% over Claude Sonnet 4.5 on Aider Polyglot. You can try it in the Poe app on all platforms and in the Poe API at: https://
poe.com/Claude-Opus-4.5. -

Why GPT 5.1 Underperforms on ARC-AGI-2 Benchmark
By
–
ARC-AGI-2 is being solved incredibly fast, but I still don’t understand why GPT 5.1 scores so low even though it’s the newest model
-

Anthropic releases Claude Opus 4.5 scoring 80.9% on SWE bench
By
–

BREAKING : Anthropic released Claude Opus 4.5 and it scores 80.9% on SWE bench verified! Where is the wall?
