AI Dynamics

Global AI News Aggregator

About

Question about RL reward-model ‘shallow’ updates vs MLE

Can you expand on "shallowly"? Do you just mean the RM's understanding of our preferences is limited, or that that RL weight updates are inherently shallow vs. MLE some other way?

→ View original post on X — @goodside