I think this is a reasonable point in theory. In practice, we don't have those model sizes. But assuming that we did, I think there's still something interesting going on with emergence, e.g., "We can't predict a performance spike at say 10B with models from 5B, 4.9B, 4.8B…"
@_jasonwei
-
Why Log-Scale X-Axis in LM Scaling Plots Matters
By
–
It's not immediately obvious why LM scaling plots use a log-scale x-axis, and as a result some people think that "emergent abilities" are not real and just an artifact of the log-scale x-axis. A quick post debunking that:
1. One reason for a log-scale x-axis is that models we -
Reasoning: The Key Differentiator Between Classical ML and Intelligence
By
–
3. The last idea is reasoning, which differentiates classical ML techniques from intelligence. Classical ML approaches need a lot of data and are black-box. Intelligent agents learn from a few examples and can do abstract reasoning.
-

Chain-of-Thought Prompting Enables Multi-Step Reasoning in Language Models
By
–
3 (cont). One way to elicit reasoning is via "chain-of-thought (CoT) prompting", which gives examples of intermediate reasoning steps in-context. CoT prompting enables large LMs to do multi-step reasoning tasks, increasing the range of tasks that LMs can do.
-
Untested Abilities and Emergent Phenomena in Scaling Large Language Models
By
–
2C. Since we haven't tested all possible abilities, we don't know the full range of abilities that have emerged in large language models.
2D. We're likely to see more emergent phenomena as we continue to scale up models (and implicit argument for more scaling). -

Unpredictable Emergence in Language Models: Key Implications
By
–
There are at least four profound implications of emergence:
2A. Emergence cannot be predicted simply by extrapolating the scaling curves from smaller models.
2B. Emergent abilities are not explicitly specified by the trainer of the language model. -

Emergence: Large Language Models Gaining Unexpected Complex Abilities
By
–
2. Emergence is a phenomenon where large language models gain abilities that are not present in smaller language models. An example of an emergent ability is doing complex math questions.
-
Three Ideas Driving the LLM Revolution: Scaling, Emergence, and Reasoning
By
–
I gave an invited lecture at New York University for @hhexiy
's class! I covered three ideas driving the LLM revolution: scaling, emergence, and reasoning. I tried to frame them in a way that reveals why large LMs are special in the history of AI. Slides: -

Scaling Laws: Model Size, Data, and Compute for LM Improvement
By
–
Key takeaways: 1. Scaling involves increasing model size, data, and compute. Scaling is challenging (cost, infra, etc), but important, since "scaling laws" tell us that scaling predictably makes LMs better.
-
Why Emergence Occurs in Larger AI Models
By
–
I don't have a great answer to "why emergence occurs", but here's a handwavy thought:
Many emergent abilities require multi-step reasoning, more knowledge, knowing how to do "tail-end" or rare tasks, and larger models are somewhat better at all of these things. With their