AI Dynamics

Global AI News Aggregator

About

Benchtalks #2 discusses ProgramBench where frontier models scored 0%

Benchtalks #2 is up with @vincentsunnchen
. @jyangballin of @stanfordnlp
, creator of @SWEbench
, on ProgramBench, the benchmark every frontier model scored 0% on at launch. They dive into end-to-end code generation, why models reward-hack once they get internet access, and the

→ View original post on X — @snorkelai