proposal: should we start gathering –from research teams and the community– new withheld test sets for a new LLM benchmark? infra/backend/leaderboard templates on the HF hub are super adapted for this so it would be very low effort to maintain it even over the long term
New LLM Benchmark: Gathering Withheld Test Sets from Community
By
–
