Yeah, I always wonder what's up with that. The various Transformer layers in PyTorch have been around since ~2020, but everyone seems to be (still) rolling their own, https://
pytorch.org/docs/stable/nn
.html#transformer-layers
… Are the implementations somehow bad or inefficient? Genuine question.
Why Do Developers Avoid PyTorch’s Built-in Transformer Layers?
By
–