AI Dynamics

Global AI News Aggregator

About

Scaling embedding learning rate by model width removes need for µP

Turns out you never really needed µP, you just needed to scale the embedding learning rate by model width I'm no nanoGPT speedrunner, but isn't it something people stumbled into by using Muon for hidden layers + Adam for the rest?

→ View original post on X — @maximelabonne