I’m thrilled to welcome @jdroege as Scale’s Chief Strategy Officer. We are obsessed with talent density at Scale—every hire should make the team better. Jason is a bar raiser. He helped launch Uber Eats, AXON, and many of his own startups. For more context on where Scale is
@alexandr_wang
-
Last US election before AGI arrival?
By
–
Potentially last US presidential election before AGI last summer Olympics before AGI too! (but that is a bit less interesting)
-
First Agent Leaderboard Published with Multiple SEAL Rankings
By
–
This is the first of many Agent Leaderboards we plan to publish. Please follow along and see our other SEAL Leaderboards:
-

ToolComp: New Tool Use Benchmark with Process Supervision
By
–
This is a new benchmark (ToolComp) that encompasses a broader range of Tool Use scenarios than prior benchmarks. Uniquely, we are further evaluating the models utilizing process supervision labels. We also split the benchmark into Enterprise and Chat use cases to differentiate
-

Models Struggle with Instruction Following and Hallucinated Information
By
–
In our error analysis, we found that most models continued to struggle with: – Final Answer Missing Information
– Hallucinated Information This indicates a widespread issue with precise instruction following, especially in prompts requiring detailed information and intermediary -

SEAL Launches Agentic Tool Use Leaderboards, OpenAI o1 Ranks
By
–
SEAL Leaderboard Update Today, we're launching 2 new Agent leaderboards:
– Agentic Tool Use (Chat)
– Agentic Tool Use Enterprise) We are also adding OpenAI o1 to the Agentic Tool Use leaderboards, ranking at: #1 on Enterprise Tool Use
#3 on Chat Tool Use (read on) -
AI Model Challenge: Submit Expert Questions for Prize
By
–
We need tough questions from human experts to push AI models to their limits. If you submit one of the best questions, we’ll give you co-authorship and a share of the prize pot. The top 50 questions will earn $5,000 each, and the next 500 will earn $500 each. All selected
-
Expert-Level AI Evaluation: Submit Your Toughest Technical Questions
By
–
If you have 5+ years in a technical field or hold/are pursuing a PhD, we want your insights! We're seeking questions that would truly impress you if an AI could solve them. Help us evaluate how close we are to achieving expert-level AI across diverse domains. Submit here:
-

Scale and CAIS Launch Humanity’s Last Exam LLM Benchmark
By
–
As LLMs get smarter, evals need to get harder.
OpenAI’s o1 has already maxed out most major benchmarks. Scale is partnering with CAIS to launch Humanity’s Last Exam: the toughest open-source benchmark for LLMs. We're putting up $500K in prizes for the best questions. (read on) -

OpenAI o1 Requires Advanced Evals for Progress Measurement
By
–
With OpenAI o1, it’s very clear we need more advanced and unsaturated evals to realistically measure progress going forward. Scale AI will have a big announcement later this week… stay tuned In the meantime, I’m back home in New Mexico, looking for eval inspiration