Validity is the degree to which evidence and theory support the interpretations of test scores for proposed uses: change the use and the question reopens.
Why it matters when the plan changes
An assessment described as validated invites the question: validated for what? The Standards for Educational and Psychological Testing, developed jointly by AERA, APA and NCME, define validity as the degree to which evidence and theory support interpretations of scores for proposed uses. A tool built to support development conversations has not thereby been validated for selecting a chief executive, even if the scores look the same.
The tension is convenience. Evidence is expensive to gather, and reusing it across purposes is tempting, especially when the new use is higher stakes and the pressure to decide is greater. Model-risk supervisors apply the same rule to their own field: the Federal Reserve's 2026 guidance warns that using a model beyond its intended purpose introduces additional uncertainty and risk. The discipline is the same whether a model prices a loan or informs an appointment.
In practice
A group rolls out a team assessment built to help teams discuss how they work together. A year later, during a restructure, the same scores are proposed as an input to deciding which managers keep their roles. The use has changed from development to consequential decisions, and the earlier evidence says nothing about whether the scores can bear that weight.
Evidence
The 2014 Standards define validity as the degree to which evidence and theory support the interpretations of test scores for proposed uses of tests (p. 11).
ETS, Validity Evidence Supporting the Interpretation and Use of TOEFL iBT Scores, citing AERA, APA and NCME (2014) (2020)The Standards are developed jointly by AERA, APA and NCME, and a new edition was announced in 2024.
Standards for Educational and Psychological Testing, Wikipedia (2026)Model-risk guidance states that using a model beyond its intended purpose introduces additional uncertainty and risk.
Board of Governors of the Federal Reserve System, OCC and FDIC, Revised Guidance on Model Risk Management (SR 26-2) (2026)
What it cannot tell you
The principle says evidence must match the use; it does not say how much evidence is enough for a given stake, and professional standards leave that judgement to test users and developers. Deciding whether a new use is genuinely new or a close variant is itself a judgement that can be argued.
Questions
From the Standards for Educational and Psychological Testing, developed jointly by AERA, APA and NCME. The 2014 edition defines validity as the degree to which evidence and theory support interpretations of scores for proposed uses. In 2024 the sponsoring organisations announced that the Standards would be updated again.
A different purpose, such as moving from development to selection; a different population, such as another country, language or level; or a different consequence, such as scores that now affect pay or employment. Any of the three can break the link between the evidence gathered and the decision now being made.
Yes. Model-risk supervisors in banking use the same logic. The Federal Reserve's revised guidance of 2026 states that using a model beyond its intended purpose introduces additional uncertainty and risk, and that extending a model's use calls for additional analysis of the new usage and its limitations.
By asking three things: validated to support which interpretation, for which population, and at what level of consequence. A claim that answers none of them is a description of research, not evidence for the buyer's decision, and the gap between the two uses should be closed before the scores carry weight.
No. Evidence for a use can weaken as populations, roles and the meaning of the work change, and new versions of an instrument need their own evidence. Validity is best treated as a standing argument that is revisited when the use, the population or the instrument itself changes.