AI Dynamics

Global AI News Aggregator

About

OrderGrad: optimization beyond the mean via order statistics

« OrderGrad: Optimization Beyond the Mean with Policy Gradient Estimation by Order Statistics » Most RL optimizes the average reward, but deployment often cares about the best sample, the worst tail, the median, CVaR, or

→ View original post on X — @askalphaxiv