Fair but it’s still actively learning to make those errors (just so it can recover from them later), which imo is still a bit weird. In an ideal world you wouldn’t, possibly process supervision is one way to get there even within RL framework. Def agree on “recovery learning”.