Skip to content

Decision-system evaluation

Decision-system evaluation is the practice of testing a whole decision process instead of only the model inside it. It asks whether the inputs, the judgement, the human review and the resulting decision each did their job, and what the outcome that followed teaches the next call.

A technically accurate signal can still be unhelpful if it arrives late, is misunderstood or changes the wrong decision, so the whole system is what gets tested.

Why it matters when the plan changes

Adding a model to a decision does not guarantee a better decision. Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone's meta-analysis of 106 experimental studies, reporting 370 effect sizes, found that on average combinations of humans and AI performed worse than the better of the two alone, with losses in tasks that involved making decisions. Testing the model by itself would miss that; only testing the combination shows it.

The tension is cost and delay. Evaluating a decision system means waiting for outcomes, recording what the reviewer changed and why, and comparing the result with the forecast, all of which is slower than scoring a model on a test set. It is also the only evaluation that tells a leader whether the process deserves the weight it is being given. A process can feel rigorous and still add nothing measurable.

In practice

A leadership team has used a quarterly execution forecast for a year. Instead of asking only whether the ratings were right, the review asks four things: were the inputs current when each call was made, did the reviewers add evidence the model lacked, did any forecast change a decision, and did the changed decisions do better than the unchanged ones.

Evidence

What it cannot tell you

Decision-system evaluation needs outcomes, which arrive slowly and are shaped by many causes besides the decision. It can show whether a process tracked reality over a set of calls; it cannot prove that any single decision was right, and the method it describes has not yet been validated this way.

How Atlas reads it

Evals are product quality control: they test the inputs, the judgement, the human review and the outcome, and each result feeds the next forecast. Release gates are meant to be tied to them, and a newer model is adopted only when it improves the relevant task without weakening privacy, traceability or consistency. Eval gates and outcome capture are among the next workflows being made repeatable, not live features.

Questions

A model evaluation asks whether the outputs are accurate on a test set. Decision-system evaluation asks whether the whole process, the inputs, the model, the human review and the decision itself, led to better outcomes. A model can score well and still fail the decision, for example by arriving after the choice was made.

Because combinations are not reliably better. A 2024 meta-analysis in Nature Human Behaviour of 106 experiments with 370 effect sizes found that, on average, human and AI combinations did worse than the better of the two alone, and that the losses concentrated in tasks that involved making decisions.

The work result the forecast named in advance, such as decision time, a milestone or a dependency delivered, read on the date it named. Outcomes defined after the fact invite the original story to be preserved. The Federal Reserve's 2026 model-risk guidance calls the equivalent step outcomes analysis.

Yes. The useful questions are whether the reviewer understood the basis for the call, added evidence the model could not see, and recorded why any adjustment was made. A review that only approves adds cost without adding judgement, and an evaluation should be able to tell the two apart.

Only after it has been scored on enough real outcomes for its intended use. Component research can support the mechanism, but it cannot substitute for validating the complete method. Until then, the honest description is a method in validation, with its evidence and its gaps both visible.