AI Dynamics

Global AI News Aggregator

About

Weight Decay vs Learning Rate Scheduling in Model Training

The norm/wd paper showed iirc that you can get exactly the same effect with an lr scheduler. I wonder if we'd get better results by carefully tuning our schedulers instead of using wd? It's something I've wondered about for years but never got around to…

→ View original post on X — @jeremyphoward