But is there evidence that they actually do that? I mean lots of companies could theoretically do that, but is it worthwhile compared to pre-training larger models (more experts) or using that compute for additional post-training?
PS: also there are two companies that have TPUs
Evidence for MoE Scaling vs Larger Pre-training Models
By
–