As our research shows, AI has a "jagged frontier" – it is good at some tasks, bad at others in unpredictable ways. Testing AI is hard. We need a collective of experts in various fields agree to test each LLM generation as a way of seeing if AI reaches expert level in that area.
