That’s almost separate… You certainly want to learn recoveries in the event of errors at test time. You just don’t want to also learn to actively create errors.
@karpathy
-
GPT-10 Input Architecture: Pixels Over Tokenization
By
–
Haha. I am afraid people interpreted my “delete tokenizer” as “use bytes directly without BPE”, the issue is you *still* need bytes encoding arbitrariness even for that! Pixels is the only way. Just like humans. It is written. If GPT-10 uses utf8 at the input I will eat a shoe.
-
Process Supervision and Error Recovery in RL Training
By
–
Fair but it’s still actively learning to make those errors (just so it can recover from them later), which imo is still a bit weird. In an ideal world you wouldn’t, possibly process supervision is one way to get there even within RL framework. Def agree on “recovery learning”.
-

Hallucinations vs Confabulation: Rethinking RNN Behavior Terminology
By
–
It’s been a decade but yes I believe I hallucinated the term in my 2015 post on unreasonable effectiveness of RNNs. I later became aware that Geoff Hinton used “confabulate”, which is often (but I think not always) a better analogue in human psychology. It’s a bit too specific,
-
Reinforcement Learning Layers in Base Model Training
By
–
I very much hope you continue working on RL! I think it's a misunderstanding that I am suggesting we need some kind of a replacement for RL. That's not accurate and I tried to clear it but did so poorly – they layer. Layer 1 was base model autocomplete.
Layer 2 was instruct -
Collaborating with Grok 5 versus competing against artificial intelligence
By
–
I’d much rather use and collaborate with Grok 5 than compete against it. Though quite similar to chess, and “in the limit” (speaking of physics!), my value add probably trends to ~zero.
-
DCLM Core Score Model Evaluation Implementation
By
–
Thank you! I'm quite happy with the core_eval.py rewrite. I wanted to evaluate my base model with the DCLM "core score" as described in their paper, but what felt like it should surely be a simple thing of ~300 lines of code actually required me to pip install and depend on a
-
90s Living Movement: Rejecting Post-2000 Technology Adoption
By
–
There is a movement I found on Instagram where people delivery choose to live in 90s, refusing all technology after 2000. Like an intermediate form of the Amish.
-

Nanochat d32 training completes with strong metrics improvement
By
–
nanochat d32, i.e. the depth 32 version that I specced for $1000, up from $100 has finished training after ~33 hours, and looks good. All the metrics go up quite a bit across pretraining, SFT and RL. CORE score of 0.31 is now well above GPT-2 at ~0.26. GSM8K went ~8% -> ~20%,
-
Ty Minix inspires LLM operating system analogy and goals
By
–
Ty MINIX is very inspiring and was exactly on my mind as well (as I've made the LLM OS analogy often). Goals