This was designed to be a very hard test for AIs, and the questions were kept private, lowering the chance they were in the training data. PhDs with access to the internet got 34% of the questions right outside their specialty, 65%-75% inside. The new Claude 3 gets 60% overall.
