“I would not trust any claims of a superior model until I see [both] ELO points on LMSys [and] private LLM evaluation from a trusted 3rd party, such as Scale AI's benchmark. The test set must be well-curated and held secret, otherwise it quickly loses potency.”
Require secret, third‑party LLM benchmarks before trusting models
By
–
