AI Dynamics

Global AI News Aggregator

About

Evidence for MoE Scaling vs Larger Pre-training Models

But is there evidence that they actually do that? I mean lots of companies could theoretically do that, but is it worthwhile compared to pre-training larger models (more experts) or using that compute for additional post-training?
PS: also there are two companies that have TPUs

→ View original post on X — @rasbt