oh cool, is it *pretraining* data? and they do the same gradient-matching thing? i wonder what the advantage is over just mixing in and doing supervised training on a few examples every now and again
By
–
oh cool, is it *pretraining* data? and they do the same gradient-matching thing? i wonder what the advantage is over just mixing in and doing supervised training on a few examples every now and again