If you have prod responses reviewing the worst ones usually finds fixable issues; try to find some crude measure of quality (flags or other behavioral indicators from users, or LLM-as-judge) to find cases where the model misunderstands the prompt.
Methods to Find and Measure Problematic AI Model Responses
By
–