Evals
Evals are a fixed set of test cases with known-good answers, run against an AI system to measure whether a change made it better or worse. They're the equivalent of a test suite for a probabilistic system.
Without evals, improving an AI feature is guesswork: someone changes a prompt, tries three examples, and declares it better. Evals replace that with a number — this change moved accuracy from 82% to 87% on 200 cases.
They don't need to be elaborate. Fifty real questions with correct answers, run automatically, catches most regressions. The discipline is building the set from real usage rather than invented examples.
Why it matters
Ask whether a proposed AI build includes an evaluation set. If not, nobody will be able to tell you whether it's getting better or worse after launch.
Related terms
More in AI Automation & Agents
Need this built rather than explained?
We publish every price we charge, and you get a quote in writing before anything starts.