AI Dynamics

Global AI News Aggregator

About

Diminishing Returns in Language Model Training Efficiency

the returns have been diminishing for a while, that’s certainly true NeoBERT is 250M params but trained on 2T tokens. objectively a crazy thing to do

→ View original post on X — @jxmnop