AI Dynamics

Global AI News Aggregator

About

LLaMA evaluation metrics concern and measurement discrepancies

It is true that accuracy on some metrics can be quite sensitive to the prompt, however this is not normal that all metrics reported for LLaMA here are systematically (and significantly) below what we measured. There may be an issue in how LLaMA was evaluated.

→ View original post on X — @guillaumelample