Summary of stable FP8 training recipe for Llama models:
1. Clip learning rate when 2nd moment estimator is outdated by coming spike – see https://
arxiv.org/abs/2304.13013
2. SmoothQuant: smooths activation outliers by migrating the quantization difficulty to weights – see
Stable FP8 Training Recipe for Llama Models
By
–
