7B model, 1M context window! Lots of results over long token windows; this looks very promising. https://
largeworldmodel.github.io RingAttention is what makes this works, summary generated by GPT: Ring Attention is a new method that reduces the memory requirements of Transformers. It
7B Model Achieves 1M Context Window with RingAttention
By
–
