Benchtalks #2 is up with @vincentsunnchen. @jyangballin of @stanfordnlp, creator of @SWEbench, on ProgramBench, the benchmark every frontier model scored 0% on at launch.
— Snorkel AI (@SnorkelAI) 3 juin 2026
They dive into end-to-end code generation, why models reward-hack once they get internet access, and the… https://t.co/WZmhUqa8ya
Benchtalks #2 is up with @vincentsunnchen
. @jyangballin of @stanfordnlp
, creator of @SWEbench
, on ProgramBench, the benchmark every frontier model scored 0% on at launch. They dive into end-to-end code generation, why models reward-hack once they get internet access, and the