What if your LLM could process long contexts without the quadratic attention bottleneck? Enter Stem: a plug-and-play sparsity module that rethinks causal information flow. It uses a token position‑decay strategy (keeping early tokens for recursive dependencies) and an
Stem: Efficient Long-Context LLM with Token Position-Decay
By
–
