AI Dynamics

Global AI News Aggregator

About

Scaling Language Models with Infinite Compute via Ensemble Distillation

As data and compute scales differently, what do we do when we have infinite compute? To answer this, this paper shows you can still scale LMs by cranking up regularization, training many small clones on the same corpus, then ensemble → distill, netting 5x data-efficiency

→ View original post on X — @askalphaxiv