Skip to content

Incremental validity

Incremental validity is the extent to which a new measure improves prediction of an outcome beyond what existing measures already provide. It is tested by adding the new measure to a model that contains the established predictors and asking how much the explained variance or accuracy rises.

A measure earns its place by what it adds to what is already known, not by how well it predicts on its own.

Why it matters when the plan changes

Most new assessments can show some relationship with performance. The harder question is whether they add anything to what an organisation already knows from a structured interview, a work sample or a track record. Frank Schmidt and John Hunter's 1998 review of 85 years of research reported that general mental ability combined with an integrity test reached a mean validity of .65, and paired combinations of that kind are how an increment is usually expressed.

The tension is that those figures have since been revised. Paul Sackett and colleagues re-examined the corrections behind them in 2022 and found mean validity estimates reduced by .10 to .20 points, with structured interviews now ranked first. Increments are sensitive to how the underlying validities are estimated, so a claimed increment is only as strong as the baseline it is measured against.

In practice

A company already hires managers through a structured interview and a reference check. A vendor proposes adding a personality inventory. The right question is not whether the inventory predicts performance, but whether it predicts anything the interview and references do not already capture, for these roles, and whether that increment justifies its cost and each candidate's time.

Evidence

What it cannot tell you

Incremental validity depends on the baseline chosen, the outcome measured and the population studied. A measure can add little over one set of predictors and a great deal over another, and an increment found in one role or country does not transfer automatically to another.

Questions

Usually with hierarchical regression: fit a model with the existing predictors, add the new measure, and see how much the explained variance or accuracy rises. The increment has to be judged against the outcome, the population and the cost, since a small statistical gain can still be expensive to collect.

Frank Schmidt and John Hunter's 1998 review reported mean validities of .65 for general mental ability plus an integrity test and .63 for ability plus a structured interview. Those figures were later revised downward, so they are best read as a ranking rather than as precise values.

Because of how range restriction was corrected. In 2022 Paul Sackett, Charlene Zhang, Christopher Berry and Filip Lievens argued the usual corrections substantially overcorrected, and found mean validity estimates reduced by .10 to .20 points. Most procedures kept their rank, with structured interviews moving to the top.

Because a vendor can truthfully report that a tool predicts performance while it adds almost nothing to what the buyer already does. The useful question is the increment over the current process for the buyer's roles, which requires evidence gathered against that baseline, not a validity figure from elsewhere.

Not according to John Hunsley and Gregory Meyer, who wrote in 2003 that most areas of applied psychology had made insufficient effort to evaluate it. They also noted the difficulty of applying group-level findings to an individual case, a limit that applies to any increment reported for a population.