Maybe you are right, we will see. However, the name GPT will definitely be gone
LLMS
-
OpenAI unlikely to merge GPT-5 and o1 before 2026
By
–
You are right, they are going to merge them. However, probably not 2025, as inference-cost for o1 is still too high. That's why I assume there will be one "regular" LLM like GPT-5 (or whatever it's called) and one reasoning model. No merge before 2026
-
Reaching GPT-4 Limit Indicates High Performance Achievement
By
–
reaching Gpt-4 limit = high performer
-
o1’s CoT: Real Thinking vs. Imitation in AI Models
By
–
Jasein Wei, @OpenAI researcher has written a great post on how CoT at o1 differs from CoT prompting. And it is clearly different. CoT prompting ultimately just imitates thinking, but doesn't actually do it. CoT at o1, on the other hand, triggers an inner monologue in the model,
-
LLMs Without Reasoning Methods Will Continue Developing
By
–
As I already wrote in my post, I assume that it only refers to LLMs that work without reasoning methods, such as CoT in o1.
In this respect, I think that we will continue to see a lot of development.
Besides, @Sama
, do you really want to leave it to Gary Marcus? @polynoamial ? -
Karpathy Explains Moravec’s Paradox and LLM Math Benchmark Challenges
By
–
Karpathy is such a goat. We learn so much from him, today about Moravecs paradox and why the new math benchmark is so difficult for LLMs.
-
Future AI Models o2 and o3 Will Deliver Significantly Better Results
By
–
I am pretty sure next models like o2 or 3 will have much much better results
-

MathFrontier Benchmark: New Challenge for LLM Capabilities
By
–
Very exciting article about the new MathFrontier benchmark, which they say LLMs are currently still failing at. The tasks are so difficult, so specialized, that perhaps PhDs from the special field can solve a task, but certainly not in general. It will be very exciting to see
-
Anthropomorphizing trap but wrong CoT outperforms non-CoT
By
–
I mostly agree — anthropomorphizing is a trap — but the choice of “Let’s think step by step” isn’t arbitrary either. Wrong CoT working better than non-CoT still says something
-
LLM Development Limits: Orion and Gemini Model Improvements
By
–
Another interesting perspective on the development of Orion. As I said, though, I believe that this refers to the fact that regular LLMs are slowly reaching their limits. That's why Google is having increasing trouble improving their Gemini models. However, since we can