And Google now running with Meta's approach of training smaller models with massive data sets. If you're curious about size, the open source RedPajamas open data set is 30T tokens. We can prob safely assume Google has access to much more.
Google adopts Meta’s approach training smaller models massive datasets
By
–
