AI Dynamics

Global AI News Aggregator

About

Transformers’ Interpolative Architecture and Limitations for Symbolic Tasks

Ironically, Transformers are even worse in that regard — mostly due to their strongly interpolative architecture prior. Multi-head-attention literally hardcodes sample interpolation in latent space. Also, the fact that recurrence is a really helpful prior for symbolic programs.

→ View original post on X — @fchollet