I only say partially since we can measure the impact of AI improvement more directly on meaningful tasks: better coding, fewer hallucinations in summarizing a complex medical case, etc. In these situations, experts can see meaningful small differences among models as they improve
Measuring AI Model Improvements Through Expert-Evaluated Tasks
By
–