We really need better benchmarks for LLMs. This paper shows that open source AIs can successfully guess the answer to standard multiple choice tests used to measure AI… even if they aren’t given the question! That suggests these tests aren’t that useful https://
arxiv.org/pdf/2402.12483
.pdf
…
Open Source AIs Gaming Standard LLM Benchmark Tests
By
–
