AI Dynamics

Global AI News Aggregator

About

Training step with refined scaffold and reward hacking guards

Each training step has the model propose a refined scaffold for a task, which it then uses to generate a solution, with reward flowing back to both stages. There are three layers of guard against reward hacking. As DeepReinforce states, the 9B variant achieves a score of 43.1 on

→ View original post on X — @testingcatalog