Ok, fair. But if you look at the TransformerEncoderLayer class, for example, I don't think it's more abstract than nn.Conv2D or nn.RNN, for example. For some reason people just like implementing Transformer layers from scratch way more than convolutional or recurrent layers .
TransformerEncoderLayer Abstraction Compared to Conv2D and RNN
By
–
