"Self-Distillation of Hidden Layers for Self-Supervised Representation Learning" This paper shows that self-supervised vision models would work much better when they predict a teacher's hidden layers across the entire visual hierarchy. As they showed that just learning from the
Self-Distillation Hidden Layers Self-Supervised Vision Models
By
–
