“Tapered Language Models” Most LMs give every layer the same MLP width, but the paper shows this is probably wasteful. Early layers seem to write more new information into the residual stream, while later layers mostly refine what is already there. So instead of making the
Tapered LMs: Early layers write more, later layers refine
By
–
