Skip to content

Validity for intended use

Validity for intended use is the principle that validity belongs to a particular interpretation and use of scores, not to a test in general. Evidence gathered for one purpose, population or level of consequence supports that use, and a new use needs its own evidence before the scores can carry weight in it.

Validity is the degree to which evidence and theory support the interpretations of test scores for proposed uses: change the use and the question reopens.

Why it matters when the plan changes

An assessment described as validated invites the question: validated for what? The Standards for Educational and Psychological Testing, developed jointly by AERA, APA and NCME, define validity as the degree to which evidence and theory support interpretations of scores for proposed uses. A tool built to support development conversations has not thereby been validated for selecting a chief executive, even if the scores look the same.

The tension is convenience. Evidence is expensive to gather, and reusing it across purposes is tempting, especially when the new use is higher stakes and the pressure to decide is greater. Model-risk supervisors apply the same rule to their own field: the Federal Reserve's 2026 guidance warns that using a model beyond its intended purpose introduces additional uncertainty and risk. The discipline is the same whether a model prices a loan or informs an appointment.

In practice

A group rolls out a team assessment built to help teams discuss how they work together. A year later, during a restructure, the same scores are proposed as an input to deciding which managers keep their roles. The use has changed from development to consequential decisions, and the earlier evidence says nothing about whether the scores can bear that weight.

Evidence

What it cannot tell you

The principle says evidence must match the use; it does not say how much evidence is enough for a given stake, and professional standards leave that judgement to test users and developers. Deciding whether a new use is genuinely new or a close variant is itself a judgement that can be argued.

Questions

From the Standards for Educational and Psychological Testing, developed jointly by AERA, APA and NCME. The 2014 edition defines validity as the degree to which evidence and theory support interpretations of scores for proposed uses. In 2024 the sponsoring organisations announced that the Standards would be updated again.

A different purpose, such as moving from development to selection; a different population, such as another country, language or level; or a different consequence, such as scores that now affect pay or employment. Any of the three can break the link between the evidence gathered and the decision now being made.

Yes. Model-risk supervisors in banking use the same logic. The Federal Reserve's revised guidance of 2026 states that using a model beyond its intended purpose introduces additional uncertainty and risk, and that extending a model's use calls for additional analysis of the new usage and its limitations.

By asking three things: validated to support which interpretation, for which population, and at what level of consequence. A claim that answers none of them is a description of research, not evidence for the buyer's decision, and the gap between the two uses should be closed before the scores carry weight.

No. Evidence for a use can weaken as populations, roles and the meaning of the work change, and new versions of an instrument need their own evidence. Validity is best treated as a standing argument that is revisited when the use, the population or the instrument itself changes.