Why do you keep doing this? I don’t understand the attempts to do childish gotcha moments. You are quoting me when GPT-4 was the best model & the labs were right. Scaling pre-training and inference compute both worked to make better models. Scaling has ALWAYS been logarithmic.
Scaling Laws in LLMs: GPT-4 and Compute Optimization
By
–