What if language models could learn to "think harder" only when they need to—allocating deep computation to challenging tokens while breezing through simple ones? We're excited to have Reza tomorrow in our AI4Science community giving a talk on Mixture-of-Recursions!
Mixture-of-Recursions: Adaptive Computation for Language Models
By
–
