Yep I think RL is misleading in that it restricts field of view. Eg like you mentioned you can imagine review/reflect doing a lot more – building tools for later use, or actively tuning the distribution for what to try next (instead of just sampling from policy independently as
Reinforcement Learning Field of View Limitations and Policy Optimization
By
–