Synthetic pretraining for sub-1B reasoning models Cool write-up from Tufa Labs (Matteo Saponati) on whether synthetic data augmentation actually helps very small (<1B) models reason better. They pretrain a 0.8B model with the Qwen3 architecture from scratch on 12B tokens of
Synthetic Pretraining Improves Reasoning in Sub-1B Models
By
–
