What if you could RL-train trillion-parameter LLMs on just 8 GPUs? Enter Orbit, an open-source framework that keeps the base model fixed and trains only a tiny BF16 adapter. It outperforms standard sync RL: 71% faster step times, 50% higher rollout throughput, 81% less train
Orbit: RL-train trillion-parameter LLMs on just 8 GPUs
By
–
