Sure — I’m not disputing it’s an advantage, or even that counterfactually it could have been the only important thing they did. But the benchmarks we see of the base model before any stolen data would be relevant (in post-training) seem to imply otherwise