TinyZero reproduction of R1-Zero
"experience the Ahah moment yourself for < $30" Given a base model, the RL finetuning can be relatively very cheap and quite accessible.
@karpathy
-

TinyZero: Affordable RL Finetuning Under $30
By
–
-
Building Diverse RL Environments for LLM Cognitive Strategy Development
By
–
For friends of open source: imo the highest leverage thing you can do is help construct a high diversity of RL environments that help elicit LLM cognitive strategies. To build a gym of sorts. This is a highly parallelizable task, which favors a large community of collaborators.
-
Move 37: L’émergence de découvertes surprenantes en IA par renforcement
By
–
"Move 37" is the word-of-day – it's when an AI, trained via the trial-and-error process of reinforcement learning, discovers actions that are new, surprising, and secretly brilliant even to expert humans. It is a magical, just slightly unnerving, emergent phenomenon only
-
RL vs SL: Demystifying Reinforcement Learning’s Core Concept
By
–
Yeah exactly. I get triggered when RL is dressed up in its full rigorous math formalism because it's gate-keeping an essentially trivial core idea. SL:
a token sequence comes from some 3rd party source (e.g. human demonstration), and you just train on it. RL:
you first sample a -
Deep Learning’s Legendary Appetite for Compute Resources
By
–
I don't have too too much to add on top of this earlier post on V3 and I think it applies to R1 too (which is the more recent, thinking equivalent). I will say that Deep Learning has a legendary ravenous appetite for compute, like no other algorithm that has ever been developed
-
Digital Agents as General-Purpose Tools Through Standard Interfaces
By
–
Projects like OpenAI’s Operator are to the digital world as Humanoid robots are to the physical world. One general setting (monitor keyboard and mouse, or human body) that can in principle gradually perform arbitrarily general tasks, via an I/O interface originally designed for… https://t.co/pZctgBa6PV
— Andrej Karpathy (@karpathy) 23 janvier 2025Projects like OpenAI’s Operator are to the digital world as Humanoid robots are to the physical world. One general setting (monitor keyboard and mouse, or human body) that can in principle gradually perform arbitrarily general tasks, via an I/O interface originally designed for
-

Jagged Intelligence: LLMs as Text Calculators
By
–
Yep I call it Jagged Intelligence. All of these favor thinking about current capability LLMs as tools, a bit more like text calculators.
-
AI Progress on Packaged Tasks vs Real-World Jobs
By
–
It’s done because it’s much easier to 1) collect, 2) evaluate, and 3) beat and make progress on. We’re going to see every task that is served neatly packaged on a platter like this improved (including those that need PhD-grade expertise). But jobs (even intern-level) that need
-
Open Weight Models and GPL-Inspired Licensing for AI
By
–
I feel that AIs will look fondly on such a gesture :). Cool idea from another comment: take inspiration from GPL and only allow this for open weight models.
-
Making Data Freely Available to LLMs for Training
By
–
Explicitly and eagerly available to LLMs, for free, for training and RAG, and start the movement.