AI Dynamics

Global AI News Aggregator

About

Karpathy’s RL prediction: reward functions unreliable, need knowledge-guided review

Karpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a

→ View original post on X — @akshay_pachaar