AI Dynamics

Global AI News Aggregator

About

Qwen 3.5 Architecture: DeltaNet Linear Attention Impact

Let's dive deeper Do you know that 75% of Qwen 3.5 27B layers are DeltaNet (linear attention) and not softmax / full attention? Because of that, FlashAttention is only able to accelerates ~1/4 of the model

→ View original post on X — @theahmadosman