Theory of mind (ToM) benchmark shows the performance of different LLMs on their ability to understand people's mental states. It would be great to understand what changed (for example, from GPT-3-davinci-002 to GPT-3.5-davinci-003) to make these cognitive properties emerge. 1/2
