It can learn shared patterns so that the individual experts don’t have to relearn the same info; ie it’s to reduce redundancy among the non-shared experts
Shared Experts Reduce Redundancy in Mixture-of-Experts Models
By
–
By
–
It can learn shared patterns so that the individual experts don’t have to relearn the same info; ie it’s to reduce redundancy among the non-shared experts