I very much hope you continue working on RL! I think it's a misunderstanding that I am suggesting we need some kind of a replacement for RL. That's not accurate and I tried to clear it but did so poorly – they layer. Layer 1 was base model autocomplete.
Layer 2 was instruct
Reinforcement Learning Layers in Base Model Training
By
–