AI Dynamics

Global AI News Aggregator

About

Gradient Surge in LLM Training: Weight Decay and Normalization

Why Gradients Rapidly Increase Near the End of Training This note investigates a sudden rise in gradient norms during the late stages of LLM training and identifies a surprising cause: the interplay between weight decay, normalization layers, and scheduled learning rate decay.

→ View original post on X — @askalphaxiv