Banger paper from NVIDIA on test-time compute for terminal agents. The finding is that you should sample several candidate shell commands, verify them before running one, and spend more on the verifier than on extra samples. With a GPT-5.6 Sol verifier choosing among 8 sampled
