i think your position comes from a fundamentally incomplete estimation of the parts of the graph that are not a homogeneous transformer — the loss function, the beam search, things like activation checkpointing that isnt simply autodiff, token fusion style tricks. caffe style
Beyond Homogeneous Transformers: Loss Functions and Inference Optimization
By
–