It's the training and inference that differ (and the weights, as a result of training), which is what matters. The architecture is only a compute substrate (a universal one, at that). "The model architecture is the same" is like saying "the codebase is still in Python and