GPT-4 is definitely a bigger leap than GPT-3, see:
@_jasonwei
-

GPT-4’s Compositional Abilities Demonstrated Through Creative Examples
By
–
This example of a pirate explaining taxes in the style of shakespeare indicates so many levels of compositionality. The qualitative experience of GPT-4 is remarkable and will unlock a world of new use cases!
-
GPT-4 represents significant progress in AI development trajectory
By
–
Since GPT 3.5 was delivered as an intermediate step, it may seem like GPT-3 to 4 was not as big, but IMO if you consider the trajectory of developments from 2020 until now, GPT-4 is quite a lot different. Super excited for the future of AI!
-
GPT-3 paradigm shift from task-specific models to few-shot prompting
By
–
Before GPT-3, we fine-tuned a task-specific neural network (e.g., BERT) for each task that we wanted to solve.
With GPT-3, there was a paradigm shift to using a single large language model for any task via few-shot prompting. -
GPT-4 Represents Bigger Leap Than GPT-3 Toward Human-Level AI
By
–
IMO GPT-4 is a bigger leap than GPT-3 was.
— Jason Wei (@_jasonwei) 14 mars 2023
– GPT-3 advanced AI from task-specific models to a single prompted model that is task-general
– GPT-4 is human-level on many hard tasks, and will signal a *societal* revolution where AI reaches every industry, starting with technology 🧵 https://t.co/tFaDxTCSXgIMO GPT-4 is a bigger leap than GPT-3 was.
– GPT-3 advanced AI from task-specific models to a single prompted model that is task-general
– GPT-4 is human-level on many hard tasks, and will signal a *societal* revolution where AI reaches every industry, starting with technology -
GPT-4 Achieves Human-Level Performance on Intellectual Tasks
By
–
Before GPT-4, AI was OK at surface-level statistical learning, but not usable for many tasks (e.g., reasoning).
With GPT-4, AI reaches human-level on many intellectual tasks—a qualitative shift in the “feel” of using the technology that will unlock a new world of use cases. -

Llama Model Enables Academic Research on Emergence
By
–
FWIW, I also think Llama is a huge step towards more research from academia on emergence
-

Investigating Llama’s Strong Performance and Data Leakage Concerns
By
–
While I’m generally not sure if data leakage is an issue, a lot of people are interested in the “why” of the surprisingly strong performance of Llama and it’d be great to rule that out with an experiment 🙂 (and big bench was not included). Llama is great, people should use it!
-

LLaMA Performance Analysis: Data Contamination and Benchmark Concerns
By
–
LLaMA has strong performance and is a great contribution. Some gentle feedback: it would be great to see a data contamination analysis, since most benchmarks evaluated are >2 years old, and BIG-Bench is omitted. So I'm curious if pre-training data contamination played a role.
-
Emergence in Language Models: Capabilities and Heuristics
By
–
New piece on emergence in language models by @JacobSteinhardt: https://bounded-regret.ghost.io/emergent-deception-optimization/#fnref7 I found the takeaways quite lucid:
– Capabilities that would lower training loss will emerge in the future
– As models scale up, simple heuristics tend to get replaced by complex ones