A repeatable test for AI quality: a set of real questions with known good answers, run against the system to score it — before and after every change. The AI equivalent of a test drive, done the same way every time.
Why it matters: without evals, "it seems better now" is vibes. With them, you can prove a change helped, catch regressions, and — critically — know your accuracy before customers do. This is where "passed its demo" and "actually works" part ways.