AI Dynamics

Global AI News Aggregator

About

LM-Evaluation-Harness Aggregation Error Analysis

The funniest part is that manually averaging your scores doesn't give you the same result as lm-evaluation-harness's aggregation (~0.01% error).

→ View original post on X — @maximelabonne