AI Dynamics

Global AI News Aggregator

About

DeBerta vs Mistral: Fine-tuning Parameters and Model Performance

Deberta does work quite well for seq cls, but Mistral should generally outperform it. However in this case the target modules do not cover enough (it should be all linear layers in the main backbone) so there's *much* fewer trainable params

→ View original post on X — @jeremyphoward