The only thing I can hope for is that OpenAI abandoned the Orion project in November 2023 and redirected all their efforts into inference models. It feels like they deliberately released this expensive, underperforming model just to tell us to chill and cut our expectations by
LLMS
-
GPT-4.5 Mistake: Outdated Dial-Up Computing Trope Criticism
By
–
One really bad mistake that bugs me is in the GPT4 vs 4.5 conversation (the one generated by 4.5), 4.5 asks "still buffering your responses like it's dial-up internet?". This is really bad because it clearly borrows tropes from early days computing, where an older computer is
-
GPT4 vs GPT4.5 Poll Results Surprise Karpathy
By
–
Okay so I didn't super expect the results of the GPT4 vs. GPT4.5 poll from earlier today , of this thread: https://
x.com/karpathy/statu
s/1895213020982472863
… Question 1: GPT4.5 is A; 56% of people prefer it.
Question 2: GPT4.5 is B; 43% of people prefer it.
Question 3: GPT4.5 is A; 35% of people -
Lightning Pod: Catching up on Gemini 2.0 with Logan
By
–
lightning pod: Catching up with @OfficialLoganK on everything Gemini 2.0!
-

GPT-4.5 Finally Available in ChatGPT for Premium Users
By
–
Took a few extra hours, but GPT-4.5 is in ChatGPT finally. If you're shilling out $200/month at least. Time to play.
-
Claude 3.7 Sonnet struggles with button color changes
By
–
Claude 3.7 sonnet when I ask it to change the color of a button pic.twitter.com/SGGDxJjylr
— Dan Shipper 📧 (@danshipper) 28 février 2025Claude 3.7 sonnet when I ask it to change the color of a button
-

Crashing Grok 3 by prompting a new efficient language
By
–

Looks like I crashed Grok 3 by prompting it to: 1. Come up with an entirely new, extremely efficient language for writing and speaking. 2. Naming this animal in said language.
-
Can generative AI perform clinical reasoning capabilities
By
–
Is generative #AI capable of clinical reasoning?
Our new piece @TheLancet @AdamRodmanMD https://
thelancet.com/journals/lance
t/article/PIIS0140-6736(25)00348-4/fulltext
… -
Grok 3 Hallucinations Compared to Perplexity Performance
By
–
I'm worried for myself too. Although my Grok 3 has been hallucinating more than Perplexity.
-

How I Use LLMs: Practical Guide to the LLM Ecosystem
By
–
New 2h11m YouTube video: How I Use LLMs This video continues my general audience series. The last one focused on how LLMs are trained, so I wanted to follow up with a more practical guide of the entire LLM ecosystem, including lots of examples of use in my own life. Chapters