I suspect the problem you're testing here might be too easy to see differences – you're getting nearly 100% either way. Would be interesting to see fp16 with dynamic loss scaling but without mixed precision to compare to there.
FP16 Training Optimization: Dynamic Loss Scaling Without Mixed Precision
By
–