AI Dynamics

Global AI News Aggregator

About

Identifying code errors in machine learning research benchmarks

Looks like the gzip paper I was enthusiastic about over-estimated its scores because of a bug in the code: it did top-2 knn instead of k=2. We should remember this as (yet another) a strong case for testing in ml code. I still like that it put a new idea in my toolbox. t.co/MyJA2BEExs

→ View original post on X — @dfintelligence