AI Dynamics

Global AI News Aggregator

About

Reinforcement Learning Field of View Limitations and Policy Optimization

Yep I think RL is misleading in that it restricts field of view. Eg like you mentioned you can imagine review/reflect doing a lot more – building tools for later use, or actively tuning the distribution for what to try next (instead of just sampling from policy independently as

→ View original post on X — @karpathy