The better the competition does, the better we (and the field) does. So I genuinely hope o4 pro or whatever comes next surprises us, positively.
@oriolvinyalsml
-

Frontier Models Defining AI Performance Pareto Frontier
By
–
Glad that our frontier models are literally defining the (Pareto) frontier across the board!
-
Team Excitement Around AI Model Development and Ecosystem
By
–
Thank you! The team is amazing, so much talent Excited to see what gets built around the model (by others & of course us!)
-

Pointwise vs Cumulative Metrics in Language Model Evaluation
By
–
Good catch! The numbers in the old table were pointwise estimates – pointwise performance is a bucketed estimate over context lengths, while the paper reports a cumulative average over context lengths. Pointwise and cumulative metrics are naturally incomparable and the pointwise
-
LLM-Powered Twitter Search Prototype Shifts Product Strategy
By
–
I thought it made you shift product strategy as the early prototype was an LLM powered twitter search
-
Elon Musk blocks X API access criticism and consequences
By
–
Wild, congrats! You should thank @elonmusk for blocking X/twitter API
-

Gemini 2.5 Pro achieves significant LiveBench performance improvements
By
–
A mere ~16 point jump on http://
livebench.ai. Such a good model! Gemini 2.5 Pro -
GPT-4 Turbo’s Record Score Jump Over Claude-1
By
–
Very kind of you to say this, but this isn't the largest score jump ever. I did some vibe coding to put the data https://
huggingface.co/spaces/lmarena
-ai/chatbot-arena-leaderboard/tree/main
… together, but it was a pain. "The largest lead was 69.30 points when GPT-4 Turbo took over from Claude-1 on November 16, 2023" -
Gemini 2.5 Pro Experimental Achieves Top Performance Benchmarks
By
–
Introducing Gemini 2.5 Pro Experimental! 🎉
— Oriol Vinyals (@OriolVinyalsML) 25 mars 2025
Our newest Gemini model has stellar performance across math and science benchmarks. It’s an incredible model for coding and complex reasoning, and it’s #1 on the @lmarena_ai leaderboard by a drastic 40 ELO margin. Only a handful of… pic.twitter.com/soiAVSD00jIntroducing Gemini 2.5 Pro Experimental! Our newest Gemini model has stellar performance across math and science benchmarks. It’s an incredible model for coding and complex reasoning, and it’s #1 on the @lmarena_ai leaderboard by a drastic 40 ELO margin. Only a handful of
-
Separate Leaderboards to Prevent Task Training Contamination
By
–
Congratulations on the release! A request: can you make sure to discourage training (in distribution) on the task, or at least separate the leaderboard into two? IMO this makes tasks like these lose a lot of their value.