Turns out you never really needed µP, you just needed to scale the embedding learning rate by model width I'm no nanoGPT speedrunner, but isn't it something people stumbled into by using Muon for hidden layers + Adam for the rest?
Scaling embedding learning rate by model width removes need for µP
By
–
