I don't think so? "Scaling laws" states that perplexity is highly predictable. For example, GPT-4's loss on some evaluations can be predicted with models of less than 1,000x compute. (
https://
arxiv.org/pdf/2303.08774
.pdf
…)
Should be a clear difference between things that can be predicted with
Scaling Laws Enable Predictable Model Performance
By
–