Karpathy's prediction about RL is coming true now!
— Akshay 🚀 (@akshay_pachaar) 19 juin 2026
He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a… https://t.co/0REApdfBUG pic.twitter.com/uAfW9yvn3m
Karpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a