It helps, but in this case GPT-4o often blurts out the wrong answer at the start but then realizes and fixes the mistake as it talks. If you grade liberally in that case, most naive prompts asking for reasoning will work fine. Structured Output is slightly tangential here.
GPT-4o blurts wrong answer, then self-corrects; structured output tangential
By
–