I am not saying you are wrong to expect that this should work, but my intuition is strongly that of course it wouldn't. I think words like "reasoning" aren't precise enough when dealing with AI – it can't mean what it does for a human, so what is the expectation?