AI Dynamics

Global AI News Aggregator

About

Parallax: Local Linear Attention matches FA 2/3 when using Muon

Parallax is a parametrized form of Local Linear Attention that drops the numerical solvers and matches FA 2/3 on decode. The most impressive part is that the architecture's benefit works with Muon but disappears under AdamW because the model learns to suppress it. It's

→ View original post on X — @maximelabonne