and finally we can compute membership inference success rate across all our models, ending up with this scaling law main takeaway: models trained on massive datasets (e.g. every LLM that comes out) can't memorize their training data there's simply not enough capacity
Scaling Laws Show LLMs Cannot Memorize Training Data
By
–
