oh yeah. I think his trick was just GradCache. and sharing negatives between GPUs isn’t trivial
GradCache Optimization for Distributed GPU Training
By
–
By
–
oh yeah. I think his trick was just GradCache. and sharing negatives between GPUs isn’t trivial