you're telling me an 8B param model was trained on fifteen trillion tokens? i didn't even know there was that much text in the world really interesting to see how scaling laws have changed best practices; GPT-3 was 175 billion params and trained on a paltry 300 billion tokens
Scaling Laws Evolution: 8B Model Trained on Fifteen Trillion Tokens
By
–
