Nice paper from @Cohere about the impact of code in the pretraining data: • Including code improves performance on non-code tasks. It would be nice to see the effect on math benchmarks. [1] • Optimal proportion is around 25% but there's not a lot of data points: what about
Code in Pretraining Data Improves Model Performance
By
–
