Let's dive deeper Do you know that 75% of Qwen 3.5 27B layers are DeltaNet (linear attention) and not softmax / full attention? Because of that, FlashAttention is only able to accelerates ~1/4 of the model
Qwen 3.5 Architecture: DeltaNet Linear Attention Impact
By
–
