question: i’m finetuning embedding models for a retrieval. i found large T5 encoders from GTR and SentenceT5 families significantly underperform BERT base how is this possible? feels it must be hyperparameters? unless i found the one NLP task ever where scaling doesn’t help
Fine-tuning Embedding Models: Why BERT Outperforms Larger T5 Encoders
By
–