AI Dynamics

Global AI News Aggregator

About

LLM Evaluation Improvements and the Challenge of Creating Good Benchmarks

Nice, a serious contender to @lmsysorg in evaluating LLMs has entered the chat. LLM evals are improving, but not so long ago their state was very bleak, with qualitative experience very often disagreeing with quantitative rankings. This is because good evals are very difficult

→ View original post on X — @karpathy