A measure earns its place by what it adds to what is already known, not by how well it predicts on its own.
Why it matters when the plan changes
Most new assessments can show some relationship with performance. The harder question is whether they add anything to what an organisation already knows from a structured interview, a work sample or a track record. Frank Schmidt and John Hunter's 1998 review of 85 years of research reported that general mental ability combined with an integrity test reached a mean validity of .65, and paired combinations of that kind are how an increment is usually expressed.
The tension is that those figures have since been revised. Paul Sackett and colleagues re-examined the corrections behind them in 2022 and found mean validity estimates reduced by .10 to .20 points, with structured interviews now ranked first. Increments are sensitive to how the underlying validities are estimated, so a claimed increment is only as strong as the baseline it is measured against.
In practice
A company already hires managers through a structured interview and a reference check. A vendor proposes adding a personality inventory. The right question is not whether the inventory predicts performance, but whether it predicts anything the interview and references do not already capture, for these roles, and whether that increment justifies its cost and each candidate's time.
Evidence
Incremental validity assesses whether a new assessment has more predictive ability than existing methods of assessment.
Incremental validity, Wikipedia (2026)Drawing on 85 years of research, general mental ability plus an integrity test had a mean validity of .65, and ability plus a structured interview .63.
Frank L. Schmidt and John E. Hunter, The Validity and Utility of Selection Methods in Personnel Psychology, Psychological Bulletin (1998)A re-analysis found mean validity estimates for selection procedures reduced by .10 to .20 points, with structured interviews ranked first.
Paul R. Sackett, Charlene Zhang, Christopher M. Berry and Filip Lievens, Revisiting Meta-Analytic Estimates of Validity in Personnel Selection, Journal of Applied Psychology (2022)Most areas of applied psychology have made too little effort to evaluate incremental validity.
John Hunsley and Gregory J. Meyer, The Incremental Validity of Psychological Testing and Assessment, Psychological Assessment (2003)
What it cannot tell you
Incremental validity depends on the baseline chosen, the outcome measured and the population studied. A measure can add little over one set of predictors and a great deal over another, and an increment found in one role or country does not transfer automatically to another.
Questions
Usually with hierarchical regression: fit a model with the existing predictors, add the new measure, and see how much the explained variance or accuracy rises. The increment has to be judged against the outcome, the population and the cost, since a small statistical gain can still be expensive to collect.
Frank Schmidt and John Hunter's 1998 review reported mean validities of .65 for general mental ability plus an integrity test and .63 for ability plus a structured interview. Those figures were later revised downward, so they are best read as a ranking rather than as precise values.
Because of how range restriction was corrected. In 2022 Paul Sackett, Charlene Zhang, Christopher Berry and Filip Lievens argued the usual corrections substantially overcorrected, and found mean validity estimates reduced by .10 to .20 points. Most procedures kept their rank, with structured interviews moving to the top.
Because a vendor can truthfully report that a tool predicts performance while it adds almost nothing to what the buyer already does. The useful question is the increment over the current process for the buyer's roles, which requires evidence gathered against that baseline, not a validity figure from elsewhere.
Not according to John Hunsley and Gregory Meyer, who wrote in 2003 that most areas of applied psychology had made insufficient effort to evaluate it. They also noted the difficulty of applying group-level findings to an individual case, a limit that applies to any increment reported for a population.