We just dropped the BTLM-3B-8K paper on arXiv! It distills our recipe for training SOTA LLMs:
– Extensively deduplicated dataset (SlimPajama)
– Hyperparameter search using muP
– Variable sequence length training + ALiBi
– Aggressive LR decay https://
arxiv.org/abs/2309.11568
BTLM-3B-8K: Distilling SOTA LLM Training Recipe
By
–
