This article shows that width should be allocated unevenly, with models wide at the beginning and end but narrow in the middle. Thus, the bottleneck forces better use of representations instead of wasting the
Variable-Width Transformers: Wide at Start and End, Narrow in the Middle
By
–
