Deberta does work quite well for seq cls, but Mistral should generally outperform it. However in this case the target modules do not cover enough (it should be all linear layers in the main backbone) so there's *much* fewer trainable params
DeBerta vs Mistral: Fine-tuning Parameters and Model Performance
By
–
