"Dynamic Linear Attention" Most long-context linear models compress tokens using fixed blocks or logarithmic schedules, but long texts are not uniform. Stable segments can be summarized, while the
By
–

"Dynamic Linear Attention" Most long-context linear models compress tokens using fixed blocks or logarithmic schedules, but long texts are not uniform. Stable segments can be summarized, while the