AI Dynamics

Global AI News Aggregator

About

10X Pretraining Compute vs Model Size: Clarification on Scaling

He said "10X pretraining compute" which doesn't mean 10x bigger. It can well be the same size but use more training tokens, longer contexts, and other algorithmic changes.

→ View original post on X — @rasbt