AI Dynamics

Global AI News Aggregator

About

GradCache Optimization for Distributed GPU Training

oh yeah. I think his trick was just GradCache. and sharing negatives between GPUs isn’t trivial

→ View original post on X — @jxmnop