AI Dynamics

Global AI News Aggregator

About

Long video reasoning with DeepSeek sparse attention and multimodal MoE

"Kwai Keye-VL-2.0 Technical Report" This paper makes long-video reasoning much more feasible by adapting DeepSeek Sparse Attention to a GQA-based multimodal MoE, reaching 256K context with only 3B active parameters. As dense attention makes hour-level context way too expensive,

→ View original post on X — @askalphaxiv