Yes, OpenAI runs GDPval, including against other models. I generally trust the results (due to team and the fact that other models, like Anthropic's was original leaders), but we also need independent evaluations for obvious reasons They report win-tie rate against human experts