AI Dynamics

Global AI News Aggregator

About

PPO training discourages repeating ‘a’ 100 times due to OOD risk

The latter, but it’s not “policy” as in something a human decided. Attempting to say “a” x100 risks going OOD and derailing into nonsense, and even if it doesn’t it can’t reliably count to exactly 100 in its head. So, it learns from PPO that trying is a bad idea.

→ View original post on X — @goodside