"OPRD: On-Policy Representation Distillation" On-policy distillation usually matches teacher and student only at the token probability level, throwing away the teacher’s hidden states. This paper moves the loss before the LM head, aligning student and teacher representations on
OPRD: On-Policy Distillation of Teacher Representations Before LM Head
By
–
