he’s pointing out we can just add the two layers I drew here and train those starting from pretrained GPT2 or something; don’t have to pretrain from scratch
Fine-tuning GPT2 with additional layers avoids pretraining scratch
By
–
By
–
he’s pointing out we can just add the two layers I drew here and train those starting from pretrained GPT2 or something; don’t have to pretrain from scratch