This paper shows people are asking a lot of medical questions of AI already, but we have little evidence of how good or bad this is. Most of the published research uses old models & compares to doctors. How do new models compare to the info people would have gotten without AI?
@emollick
-
Blood Test Predicts Alzheimer’s Disease Up to 25 Years in Advance
By
–
There is a sort of file drawer bias: AI benchmarks that don’t meaningfully benchmark performance are dropped, but mostly because they are either 0 or 100. The whole point of benchmarks is to measure something about AI performance. Verisimilitude is a different matter, though.
-
Evidence for Mammograms and New Papers Beyond Current Research
By
–
Having done a lot of work on measuring performance, I don’t think this is very common. There are very few tasks that AI can do where there is not an upward trajectory. You can find tasks that no AI can do, or where there is saturation, but otherwise you get improvement over time.
-
AI Benefits Breast Cancer Screening: New Reports and Editorial
By
–
This is far from an exhaustive list – look at my past tweets to find dozens more academics doing interesting work on the topic
-

Global Health Crisis: 5 Million Deaths Annually Due to Lack of Physical Activity
By
–
One thing thing about AI, for better and worse, is that "everything around me is somebody's life work" is no longer a true assumption going forward.
-
Multivitamin Slows Epigenetic Aging in Randomized Trial, Cocoa Extract Ineffective
By
–
And, we have not seen Mythos (or whatever OpenAI and Google are releasing)
-
Critique of Utah’s AI Prescriptions Based on Company Preprint and Author’s Work
By
–
Oh god, please not "load bearing" – prompt better.
-
Tripling Cancer Immunotherapy Benefit with FMT
By
–
A major lesson to take away from Opus 4.7 is that, while there is a lot of arguments about implementation choices and personality, models keep improving measurably on economically important tasks with each release (it has been two months since Opus 4.6), with no signs of slowdown
-
Questioning “Breaking News” Label for Recently Published Research Paper
By
–
I don't expect reviewers to comply with this at all, and to avoid disclosing their AI use as a result, and I don't think there is actual risks. Better to clearly specify the risk and expectations.
-
Coverage of an Interesting Research Paper on AI
By
–
Yes, OpenAI runs GDPval, including against other models. I generally trust the results (due to team and the fact that other models, like Anthropic's was original leaders), but we also need independent evaluations for obvious reasons They report win-tie rate against human experts
