I haven't seen that yet. I think the standard is to prune all linear layers, this includes the QVK of the attention layers and the FC layers (except the embedding and output layers of course).
Standard Practice for Pruning Linear Layers in Neural Networks
By
–