Why do Mixture-of-Experts models keep re-computing the same expert choices for similar inputs? Researchers from Mashang Consumer Finance, Nanjing University, and Alibaba Group introduce RMS-MoE: they add a Co-Activation Memory that remembers which expert teams worked best for
