In general, we've found that TPUs are quite effective at inference, especially for mixture of expert models and handling attention for transformer models with long context. The high speed multidimensional torus interconnects are quite fast and effective. I'm not sure I have a
TPUs excel at inference for mixture of experts transformer models
By
–