The norm/wd paper showed iirc that you can get exactly the same effect with an lr scheduler. I wonder if we'd get better results by carefully tuning our schedulers instead of using wd? It's something I've wondered about for years but never got around to…
Weight Decay vs Learning Rate Scheduling in Model Training
By
–