AI Dynamics

Global AI News Aggregator

About

Largest LLM-as-Judge audit shows exact-match overstates skill

The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, JudgeBench, and RewardBench. Findings: Validating a judge with exact-match agreement overstates its skill, because exact match does not

→ View original post on X — @dair_ai