A public, standardized eval used to compare models against each other — the league tables in AI announcements.
Why it matters: benchmark scores tell you a model is generally capable; they tell you nothing about your task on your data. Never pick a model on benchmarks alone — run your own golden set.
Related: Eval