Test-time reasoning models often converge too early. Achieving broader reasoning coverage requires longer sequences, yet the probability of sampling such sequences decays exponentially during autoregressive generation. The authors call this the "Shallow Exploration Trap." This
Shallow Exploration Trap in Test-Time Reasoning Models
By
–
