You mention our "custom prompt" like if there was an official way of prompting (yours?). Most benchmarks were created before this concept of LLM eval with prompting even exists, and for many of them there is no official prompt or way to evaluate them with LLMs.