“A Bitter Lesson for Data Filtering” A common consensus is that LLMs need carefully filtered web data, because noisy data can easily hurts in small compute regimes. But this paper shows that with a large enough model size and training, the best filter is no filter. Raw Common
