Yes, OpenAI runs GDPval, including against other models. I generally trust the results (due to team and the fact that other models, like Anthropic's was original leaders), but we also need independent evaluations for obvious reasons They report win-tie rate against human experts
RESEARCH
-
AI Researchers No Longer Need Math or Coding Skills
By
–
You no longer have to be good at math or coding to be a top AI researcher. AI does them for you.
-

Editorial on Advancing Women’s Health Prevention via Trusted Touchpoints
By
–
One of the premier journals in my field… I think there are very valid reasons to set rules on AI in peer review (including disclosure), but the idea that all AI models steal your data is very 2023. Require people to use enterprise accounts or models with training turned off.
-
AI in Mammography: New Report Calls for Integration to Prevent Breast Cancer and Heart Disease
By
–
Anyhow, I realize nobody cares and all the AI labs have started presenting their GDPval-AA score, but it is an incredibly gameable output with low face validity and we really need trustworthy measures of AI ability.
-
AI in Mammography for Cardiovascular Risk Assessment
By
–
GDPval is one of the most important benchmarks of AI ability because it is based on human expertise. It compares expert human performance to AI performance using expert human judges who spend an average of an hour evaluating each answer. It also has holdout questions that are not
-
SpaceX AI 1T Model Training Nears Completion With Major Improvements
By
–
This is still 0.5T, but a more recent training checkpoint. 1T model is ~5 days away from finishing initial training. Will be a major step change improvement in coding, long context and skills. The SpaceXAI model factory is finally working. Should be an improved base model
-
Scaling Law of P-Hacking: AI Enables Absurd Proof Generation
By
–
Scaling law of p-hacking: the more AI you use, the more absurd things you can prove.
-
Harm to Others vs Self-Harm in AI Ethics Distinction
By
–
This obfuscates the important difference between creating harm to someone vs. creating harm to yourself (by making the performance of the thing you are mutilating worse)
-

AI Model Distillation: How Stronger Models Train Weaker Ones
By
–
Everyone is accusing everyone of “stealing AI”
But almost nobody is explaining what’s actually happening. Distillation. → Query a stronger model at scale
→ Collect outputs (reasoning, code, decisions)
→ Train your own to imitate it No weights. Just behavior. This worked in -
DeepMind Leads AI for Science Among Top Labs
By
–
It would seem that DeepMind is the only top AI lab left that's serious about pushing AI for science now.