Skip to content

Confidence level

A confidence level states how much reliance a judgement can carry, attached to the judgement rather than left for the reader to infer. It answers a different question from whether the judgement is correct: it says how much weight to put on it while deciding.

Confidence tells a reader how much of the decision to rest on a finding, which is the part they have to guess when it is missing.

Why it matters when the plan changes

A reader given a finding without a confidence has to estimate one, and they estimate it from tone, seniority and how well the finding fits what they already believe. Those are the three worst available inputs. Stating confidence moves the estimate from the reader's impression to the author's evidence, and it turns the statement into a claim that can later be checked against what happened, in the way a probabilistic prediction can be scored for how well its stated confidence matched the outcome.

The tension is that confidence has to decay with distance and rarely does. Predictability has a horizon, so a judgement about this quarter and one about three years out cannot carry the same weight even when the reasoning is equally good. Forecasting itself is the practice of making predictions that are later compared with what actually happened, which is why a confidence that does not fall with the forecast horizon is not measuring anything real.

In practice

Two findings appear in the same document. One rests on current decision records and observed patterns; the other on a single interview and a document predating the restructure. Without stated confidence they read identically. With it, the leadership team knows which one to act on and which one to go and check.

Evidence

  • Probabilistic judgements can be scored for how well stated confidence matches outcomes, which is what makes a confidence level a claim rather than a decoration.

    Brier score, Wikipedia (2026)
  • Forecasting is the practice of making predictions that are later compared with what actually happened.

    Forecasting, Wikipedia (2026)

What it cannot tell you

A confidence level does not say whether a judgement is correct; it says how much weight to place on it while deciding. A high confidence attached to a flawed method still returns a wrong answer with force. The level is only as sound as the evidence behind it, and it says nothing about what that evidence actually was.

Questions

Established, where component research and governance already exist. Supported, where the reasoning is coherent enough to test. In validation, which covers the Atlas forecast itself. Direction, for the learning layer across contexts. Each says what kind of evidence is standing behind the claim.

Percentages imply a frequency that can be checked, which is appropriate for a scored forecast and misleading for a claim about a mechanism. Using them where no scoring exists borrows the authority of measurement without the measurement, which is the failure the levels are designed to avoid.

It should fall. Predictability has a finite horizon, so a judgement about this quarter carries more weight than one about three years out even when the reasoning behind both is equally careful. Confidence that stays flat across horizons is not tracking anything real.

Low confidence means a judgement was made on thin evidence and can still be acted on cautiously. Blind means no judgement could be made at all. Collapsing the two would either hide a real gap or discard a usable finding, depending on which way it was collapsed.

The author of the judgement, from what the evidence supports rather than from how certain they feel. That distinction is the whole discipline, and it is why confidence is recorded alongside the evidence and the counter-evidence rather than stated separately as a summary.