The funniest part is that manually averaging your scores doesn't give you the same result as lm-evaluation-harness's aggregation (~0.01% error).
By
–
The funniest part is that manually averaging your scores doesn't give you the same result as lm-evaluation-harness's aggregation (~0.01% error).