He said "10X pretraining compute" which doesn't mean 10x bigger. It can well be the same size but use more training tokens, longer contexts, and other algorithmic changes.
10X Pretraining Compute vs Model Size: Clarification on Scaling
By
–
By
–
He said "10X pretraining compute" which doesn't mean 10x bigger. It can well be the same size but use more training tokens, longer contexts, and other algorithmic changes.