AI Dynamics

Global AI News Aggregator

About

Falcon vs LLaMa: Evaluation methodology comparison issues

Cool though at the moment the comparison with Falcon doesn’t make any sense since the numbers come from two different evaluation setups (EleutherAI harness for Falcon on the leaderboard and Francis setup for LLaMa in this thread).

→ View original post on X — @thom_wolf