AI Dynamics

Global AI News Aggregator

About

Standard Practice for Pruning Linear Layers in Neural Networks

I haven't seen that yet. I think the standard is to prune all linear layers, this includes the QVK of the attention layers and the FC layers (except the embedding and output layers of course).

→ View original post on X — @rasbt