Our research shows that smaller foundation models that are fine-tuned on domain-specific datasets can outperform larger foundation models. We show that a GPT-NeoX 1.4B model that is fine-tuned for 2,000 training steps can perform just as well as the out-of-the-box GPT-J 6B model.
Fine-tuned Smaller Models Outperform Larger Foundation Models
By
–
