It does that because all of its training data in the last, post-training stage are of the form [question -> authoritative sounding solution], where the solutions are written by humans. The LLMs just imitate the form/style of that training data.
@karpathy
-

Jagged Intelligence: LLMs Excel and Fail Unpredictably
By
–
Jagged Intelligence The word I came up with to describe the (strange, unintuitive) fact that state of the art LLMs can both perform extremely impressive tasks (e.g. solve complex math problems) while simultaneously struggle with some very dumb problems. E.g. example from two
-

Investing in Distributed Creator Intelligence Over Concentrated Resources
By
–
I'd be a lot more inclined to invest $10M into 2000 creators. The distributed intelligence and creativity of the crowd feels underutilized.
-

LLMs Reaching LHC-Level Complexity in Infrastructure and Development
By
–
LLMs as an artifact are trending to the complexity of something like the LHC. This is clear when you look at the datacenter computronium build out but it's a lot more than that – a large chunk is digital and much harder to see/appreciate, it's just a bunch of people on a laptop.
-
Karpathy Plans Multiple Llama 3.1 Fine-tuning Projects
By
–
Yep I’d like to do many Llama 3.1 finetunes, coming up.
-

App Store Early Years Were Gimmicky, AI Will Take Time
By
–
My opinion on this has changed at a recent Sequoia event where they compared to iOS. The first ~3 years of the App Store were all kinds of gimmicky apps. I think it just takes a while to process a new thing, figure out what it is and isn't and package it into products. Image:
-
Meta’s Llama 3.1 405B: First Open Frontier LLM Available to All
By
–
Huge congrats to @AIatMeta on the Llama 3.1 release!
Few notes: Today, with the 405B model release, is the first time that a frontier-capability LLM is available to everyone to work with and build on. The model appears to be GPT-4 / Claude 3.5 Sonnet grade and the weights are -
AGI Speed Makes AI Interaction Instantly Pleasing and Natural
By
–
This is so cool. Feeling the AGI – you just talk to your computer and it does stuff, instantly. Speed really makes AI so much more pleasing.
-

New favorite LLM test reveals inconsistent performance across SOTA models
By
–
Wow, this has just become my favorite LLM test. I missed that this doesn't work but it really doesn't, even for SOTA LLMs. Seems to be a bit hit and miss, e.g. with GPT4o which failed 1/3 times, Claude failed 3/3 times.
