AI Dynamics

Global AI News Aggregator

About

Memory Sparse Attention Scales Models to 100M Token Context

Scaling Attention to 100M context!? Memory Sparse Attention introduces an idea where instead of rereading an entire 100M-token entry, it learns to jump straight into the relevant memories and reason from them end-to-end. More specifically, it first encodes documents into

→ View original post on X — @askalphaxiv