We need more real-world benchmarks for models. Saying that a model scored "87.4% on the MMLU AQuA-RAT" is useless for anyone who's not a researcher. How about we test for: -ER diagnosis (accuracy + time)
-Radiology reads (scan → diagnosis)
-ICU management (decisions →
Proposing real-world benchmarks for medical AI models
By
–