"Kwai Keye-VL-2.0 Technical Report" This paper makes long-video reasoning much more feasible by adapting DeepSeek Sparse Attention to a GQA-based multimodal MoE, reaching 256K context with only 3B active parameters. As dense attention makes hour-level context way too expensive,
Long video reasoning with DeepSeek sparse attention and multimodal MoE
By
–
