AI Dynamics

Global AI News Aggregator

About

µP embedding LR rule correct under AdamW, explains most benefits

To clarify, this paper basically says: under AdamW, µP's embedding LR rule (constant) is essentially right and explains most of µP's benefit. Last year, Hayou et al. found that µP's embedding LR rule is wrong for realistic LLM vocab sizes. They found that the optimal embedding

→ View original post on X — @maximelabonne