This is good news, but the fact that Composer 2.5 comes out on top over GLM 5.2 here should trigger a serious re-evaluation of the entire benchmark. It's nowhere near in practice except basic tasks… and that undermines trust in Cursor's entire model evaluation / publicity.
Composer 2.5 beating GLM 5.2 triggers benchmark re-evaluation
By
–
