The key challenge of training any non-English model is data. To build a dataset large & rich enough for Arabic, Jais tapped into various open and curated sources totally 55B tokens. The model trained on 2 epochs of this data.
Building Arabic Language Models: Data Collection and Training Strategy
By
–
