AI Dynamics

Global AI News Aggregator

About

Harness Evaluation Framework for Large Language Models

18/ Harness: Now we finally turn to the EleutherAI Harness implementation (as of January 2023) which was used to compute the numbers for the Open LLM Leaderboard. Here is yet another way to compute a score for the model on the very same evaluation dataset! Let's take a look:

→ View original post on X — @thom_wolf