Introducing long-context transformer using mini sequences. It is a simple and effective method for highly efficient and accurate LLM training with extremely long sequences. Our research demonstrates that the Llama3-8B model can be trained with context lengths up to 60k tokens on
Long-Context Transformers: Training Llama3 with 60k Token Sequences
By
–
