Introducing AltUp, a method that takes advantage of increasing scale in Transformer networks w/out increasing the computation cost — it’s easy to implement, widely applicable to Transformer architectures, and requires minimal hyperparameter tuning. More → https://
goo.gle/475j3iW
AltUp: Efficient Transformer Scaling Without Increased Computation
By
–
