Yes, this explains how it works. The primary method for training this model is behavior cloning, and I explain what it is along with reinforcement learning (RL) and its limitations. And I don’t mention bias at all, sorry.
Explanation of Behavior Cloning and Reinforcement Learning
By
–