we train all of our models until they "saturate" which usually happens around 1M steps using a very large batch size models memorize the same amount, regardless of training datasize meaning they have fixed capacity and instead "spread it thinner" when trained on more examples
Model Memorization Fixed Capacity Training Saturation Analysis
By
–
