AI Dynamics

Global AI News Aggregator

About

FP16 Training Optimization: Dynamic Loss Scaling Without Mixed Precision

I suspect the problem you're testing here might be too easy to see differences – you're getting nearly 100% either way. Would be interesting to see fp16 with dynamic loss scaling but without mixed precision to compare to there.

→ View original post on X — @jeremyphoward