Totally. Also, expert-created private evaluations that are impossible to overfit like SEAL (disclaimer, this is our work at Scale)
@alexandr_wang
-
AI Evaluation Ecosystem Needs Better Benchmark Overfitting Detection
By
–
The whole Reflection-70B debacle points the the desperate need for a better AI evaluation ecosystem. It needs to be extremely easy to adjudicate:
(1) is the model overfit to benchmarks
(2) is the model truly unique (i.e. not a wrapper or thin fine-tune) https://
x.com/shinboson/stat
/shinboson/status/1832933747529834747
… -

Human Red-Teamers Outperform Automated Methods in Multi-Turn Jailbreaks
By
–
ICYMI: New research from SEAL that demonstrates human red-teamers massively outperform automated methods over multiple turns. We also release MHJ, a dataset of multi-turn jailbreaks to help enable further research into multiturn red teaming
-

PlanSearch: New SOTA Test-Time Compute Method for Code
By
–
New SOTA test-time compute result from Scale SEAL We are releasing a new SOTA test-time compute method called PlanSearch. It meaningfully outperforms existing approaches on LiveCodeBench via a new diversity-based search method See more about our SEAL open research below: https://
x.com/hughbzhang/sta
/hughbzhang/status/1832079839840575705
… -
New Math Benchmark Developed to Evaluate AI Models
By
–
Given that our Math leaderboard (GSM1K) is now relatively saturated (most models score >90), we are working on producing a new Math benchmark to properly discern between models.
-
Grok 2 Updates Coming in Upcoming Weeks
By
–
Expect more updates, including the addition of Grok 2, in the coming weeks!
-

Claude 3.5 Sonnet and Llama 4.1 405B Excel in Instruction Following
By
–
INSTRUCTION FOLLOWING: Claude 3.5 Sonnet and Llama 4.1 405B Instruct stand out as models with BOTH: – very high (>0.9) Main Request fulfillment
– very high (>0.7) Constraint fulfillment Also, all Claude models take a hit for Structural Clarity in writing style -

Models Show Weaker Instruction Following Performance in Spanish
By
–
MULTILINGUAL (SPANISH) Lastly, one interesting pattern we noticed is that every model seemed to perform worse at Instruction Following (Main Request fulfillment and Constraint fulfillment) in Spanish versus English. This implies there's still meaningful headroom for models to
-

GPT-4o leads code correctness; Claude excels prompt adherence
By
–
Some insights from the leaderboards: CODING We notice that GPT-4o (August 2024) performs the best on Code Correctness among all the models, but Claude outdoes it on Prompt Adherence