Current RL fine-tuning for LLMs suffers from expensive training, critic networks or advantage estimates, and inefficient exploration. Enter A*-PO: A two-stage policy optimization method that flips the standard approach — slashing compute and memory costs compared to
A*-PO: Efficient Two-Stage Policy Optimization for LLMs
By
–
