If we could train all our models from scratch on publicly scrutinized data that’d be great an I hope it happens! But that’s not the world right now. Even “open source” models don’t distribute their pre-train data — weights are closer to binaries than source code.
Data Scrutiny and Open Source AI Model Transparency Challenges
By
–