An eval turns a question of impression into a repeatable test, and a result only counts if the same test ran before the change.
Why it matters when the plan changes
Systems built on models change constantly: a new model version, a revised prompt, a different data source. Without a fixed test, each change is judged by impression, and a regression can ship because the new output reads well. Evals make quality comparable across versions. OpenAI's public evals repository, for example, describes itself as a framework for evaluating language models and an open-source registry of benchmarks.
The tension is that an eval measures only what it was written to measure. A model can pass every case in a suite and still fail in the setting that matters, because the suite missed it or because conditions moved. Supervisors of model risk address this with outcomes analysis, comparing model outputs with corresponding real-world outcomes, alongside the tests run before a model is first used. Passing a suite is evidence about the suite first.
In practice
A team swaps the language model behind an internal assistant for a cheaper one. Before the switch it runs the same two hundred test questions through both, scored for correct sourcing and for declining to answer when evidence is missing. The cheaper model matches on accuracy but cites sources less reliably, so the switch waits until that gap is closed.
Evidence
OpenAI's evals project describes itself as a framework for evaluating language models and systems, with an open-source registry of benchmarks.
OpenAI, Evals repository (2023)NIST's AI Risk Management Framework, released on 26 January 2023, is voluntary and covers the design, development, use and evaluation of AI systems.
National Institute of Standards and Technology, AI Risk Management Framework (2023)Model-risk guidance defines outcomes analysis as comparing model outputs with corresponding real-world outcomes.
Board of Governors of the Federal Reserve System, OCC and FDIC, Revised Guidance on Model Risk Management (SR 26-2) (2026)
What it cannot tell you
An eval result describes performance on its own test cases at the time it ran. It does not show performance on cases the suite omits, it can be gamed once the suite becomes a target, and it says nothing about whether the thing measured was the thing that mattered.
Questions
A benchmark is usually a public, shared test used to compare systems across organisations. An eval is any repeatable test a team runs on its own system, often built from its own cases. Many evals include benchmark items, but the ones that matter most are written for the specific use.
Because a newer model is not automatically better at a particular task. A change of model, prompt or data source can improve general scores and weaken the one behaviour a product depends on. Running the same eval before and after each change is how that trade is noticed before users notice it.
There are frameworks rather than one standard. NIST's AI Risk Management Framework, released on 26 January 2023, is voluntary and covers evaluation alongside design, development and use. In banking, the Federal Reserve's model-risk guidance of 2026 expects outcomes analysis that compares model outputs with real-world outcomes.
Yes. Once a score becomes the target, systems and teams optimise for the test cases rather than the task. Rotating cases, holding some back and checking results against real outcomes all reduce the risk, which is why model-risk guidance from the Federal Reserve in 2026 pairs testing with outcomes analysis.
It reflects a real situation the system will meet, has an answer or scoring rule agreed before the run, and would fail if the behaviour it guards against returned. A case that cannot fail is not a test. The best suites grow from real errors, each one turned into a case that must pass from then on.