Ya. It’s GSM1k under the hood. We’ll have a harder one soon.
@alexandr_wang
-
GSM-1k Benchmark Limitations and Upcoming Harder Math Evaluation
By
–
GSM-1k wasn’t really designed to distinguish between top models, more to detect overfitting. we will fix this for the next round with a harder math eval!
-
Flash Praised as Excellent Small Language Model
By
–
Yes, I’ve said this before, but Flash is a crazy good small model
-
Periodic AI Model Evaluation Updates and Benchmarking Strategy
By
–
5/ We plan to periodically update these evaluations with: new evaluation sets, to ensure fresh, non-saturated results new models, as they become available Please let us know what you think!
-
Third-party AI evaluations and overfitting prevention strategies
By
–
4/ While LMSYS and other efforts in the community are awesome, we still think there's a lot to be desired in 3rd party evaluations. One of our design principles is to produce evals that are impossible to overfit. As we saw with our prior GSM1k research, we think it's critical
-

Leading AI Models Evaluated: Coding, Math, Multilingual Performance
By
–
3/ We eval'd many of the leading models: – GPT-4o
– GPT-4 Turbo
– Claude 3 Opus
– Gemini 1.5 Pro
– Gemini 1.5 Flash
– Llama3
– Mistral Large On Coding, Math, Instruction Following, and Multilinguality (Spanish). See leaderboard results below. -
Third-party AI evaluations critical to ecosystem development
By
–
2/ Evaluations are a critical component of the AI ecosystem. Evals are incentives for researchers, and our evaluations set the goals for how we aim to improve our models. Trusted 3rd party evals are a missing part of the whole ecosystem, which is why @scale_AI built these.
-

Scale Launches SEAL Leaderboards for Frontier Model Evaluation
By
–
1/ We are launching SEAL Leaderboards—private, expert evaluations of leading frontier models. Our design principles:
Private + Unexploitable. No overfitting on evals!
Domain Expert Evals
Continuously Updated w/new Data and Models Read more in http://
scale.com/leaderboard -
Scale AI Bullish on San Francisco Long-Term Future
By
–
very bullish on San Francisco long-term @scale_AI is a proud SF company, and has been since 2016!