2/5 When sequence lengths vary a lot, as in RL training, padding can consume a large fraction of the tokens. This wastes compute and memory and drives up training time and cost.
Padding Waste in RL Training: Impact on Compute and Cost
By
–

By
–

2/5 When sequence lengths vary a lot, as in RL training, padding can consume a large fraction of the tokens. This wastes compute and memory and drives up training time and cost.