Skip to content

Predictive validity

Predictive validity is the extent to which a score predicts a later outcome that matters. It is established by comparing the score with the outcome after the fact, for a specific use and a specific population, and it does not carry over to a different use or a different group.

Validity is a property of a use, not of an instrument; evidence for one purpose does not transfer to another.

Why it matters when the plan changes

An instrument is frequently described as validated with no statement of what it was validated to predict, for whom, or how well. Those three qualifiers are the whole claim: predictive validity is the extent to which a score predicts a specific later outcome, for a specific population. Without them, the word does reputational work rather than evidential work, and a buyer has no way to tell which they are being sold. The Uniform Guidelines on Employee Selection Procedures treat this specificity as a requirement for decisions that affect employment, not a nicety.

The tension is timing. Predictive validity can only be established after the outcomes exist, so a new instrument cannot have it yet, by definition. Standards for reliability and use-specific validity evidence, required before high-stakes application, make the same point: construction and mechanism are not themselves evidence of prediction. The honest position in that interval is to describe the mechanism, the construction and the validation plan, and to decline the claim until the outcomes are in.

In practice

A vendor cites a validity figure from a study on graduate selection in one country as support for using the same instrument on executive appointments in another. The figure is real and it is evidence for a different use and a different population, which makes it irrelevant to the decision in front of the buyer.

Evidence

What it cannot tell you

Predictive validity says a score correlates with a specific later outcome for a specific population; it says nothing about why, and nothing about whether the same score predicts a different outcome or holds for a different group. A high figure for one use gives no basis for judging performance on a use that has not yet been tested.

Questions

Three qualifiers: what outcome was predicted, for which population, and how strongly. A claim missing any of the three cannot be evaluated. The Uniform Guidelines on Employee Selection Procedures, issued in 1978, treat validity evidence for the specific use as a requirement, which is exactly why the use has to be named.

No. Evidence gathered for one purpose or one population does not carry to another. An instrument validated for graduate selection has not thereby been validated for executive appointment, and treating the two as equivalent is the most common way the word is misused.

By recording the scores, waiting for the outcomes, and comparing the two. That sequence cannot be shortened, which means a new instrument does not have predictive validity yet regardless of how well it is constructed or how strong the mechanism behind it is.

The mechanism, the construction, the reliability work and the validation plan, each described as what it is. The Wikipedia entry on predictive validity (2026) defines the term but says nothing about strength of prediction; a definition is not evidence. Research can support a mechanism; only intended-use evidence supports a claim.

Validity asks whether the score predicts the outcome at all. Calibration asks whether stated confidence matches observed frequency: whether things called seventy percent likely happen about seventy percent of the time. A judgement can be valid and badly calibrated, or well calibrated and uninformative.