Global AI News Aggregator
About
By
–
same – for models of that size, not sure why you actually need distributed training either.
→ View original post on X — @nathanbenaich