We are in the era of $5 Uber rides anywhere across San Francisco but for LLMs weee
@karpathy
-
Voice Quality Regression in AI Assistant After Updates
By
–
They iterated on it a bit, e.g. custom instructions and the ability to join the podcast, but I think overall agree. I actually used it again this morning after a while and it felt a bit regressed, even? The woman's voice esp sounds slightly more dead / less interested, or
-
Agency More Powerful and Scarce Than Intelligence
By
–
Agency > Intelligence I had this intuitively wrong for decades, I think due to a pervasive cultural veneration of intelligence, various entertainment/media, obsession with IQ etc. Agency is significantly more powerful and significantly more scarce. Are you hiring for agency? Are
-
Fine-tuning’s Surprising Power: Knowledge Retention Through Style Change
By
–
Wow it really has been that long 😐
The big thing I didn’t realize is that an assistant was just a finetune away. That is the surprising thing I was really missing. I still find it surprising today, that you can just change the style so dramatically but retain the knowledge. -
Dataset difficulty balance: avoiding trivial and impossible problems
By
–
Great question it doesn’t. You hope there are always enough problems in your dataset that are *just right* hard – not trivial and not impossible. If not, it won’t work.
-
Task Evals vs AGI: The Gap Between Simple Tests and Autonomous Agents
By
–
Let's keep in mind these are still super simple "task" evals. Little queries served on a platter, even if increasingly difficult. Which are super helpful, but when people talk about AGI they usually have an autonomous agent swarm in mind performing long-running jobs across
-

Reproducing Software Build Issues: Testing Consistency Analysis
By
–
Hah yeah I can reproduce it and got something similar 5/5 attempts. I wonder what build they sent him earlier
-
Why Humor in LLMs Remains Difficult to Achieve Through Post-Training
By
–
Great question right? I'd love to know, I don't think I fully understand this either. But considering that noone has (to my knowledge) figured out a way to post-train an LLM to be funny, I am prepared to believe humor is really difficult and requires more underlying capability?
-

Grok 3 Early Access: State-of-the-Art Thinking Model Review
By
–
I was given early access to Grok 3 earlier today, making me I think one of the first few who could run a quick vibe check. Thinking First, Grok 3 clearly has an around state of the art thinking model ("Think" button) and did great out of the box on my Settler's of Catan