Why Gradients Rapidly Increase Near the End of Training This note investigates a sudden rise in gradient norms during the late stages of LLM training and identifies a surprising cause: the interplay between weight decay, normalization layers, and scheduled learning rate decay.
Gradient Surge in LLM Training: Weight Decay and Normalization
By
–
