An agent estimating probabilities on how to act is mathematically estimating expectations over sets. So it can learn expected returns or value signals. But, again, that is NOT how rewards are used in modern AI. In AI rewards are often defined and created by humans. This, at its
On agents estimating expectations and the role of rewards
By
–