A big issue I see with AI systems is that people aren't spending enough time evaluating their evaluation pipeline. 1. Most teams use more than one metrics (3-7 metrics in general) to evaluate their applications, which is a good practice. However, very few are measuring the
Evaluating AI Systems: The Overlooked Evaluation Pipeline Problem
By
–