the returns have been diminishing for a while, that’s certainly true NeoBERT is 250M params but trained on 2T tokens. objectively a crazy thing to do
Diminishing Returns in Language Model Training Efficiency
By
–
By
–
the returns have been diminishing for a while, that’s certainly true NeoBERT is 250M params but trained on 2T tokens. objectively a crazy thing to do