Sorry, but I think you’re misinterpreting the paper: The 62% is an ensemble of TTT + BARC on the public validation set and doesn’t imply any score on the harder semi-private v1. It shouldn’t be on that plot.
Clarifying reported model ensemble scores on validation vs semi-private set
By
–