Samuel Messick described construct validity as one concept with six distinguishable aspects, from the content of a test to the consequences of using its scores.
Why it matters when the plan changes
An instrument can be reliable and still measure the wrong thing. A scale labelled strategic thinking may mostly capture verbal fluency or confidence, and every later finding inherits that error. Construct validity asks whether the scores reflect the intended concept and stay distinct from its neighbours, and it comes before any claim about prediction. Samuel Messick's 1994 report set out six aspects of the evidence: content, substantive, structural, generalisability, external and consequential.
The tension is that the evidence never closes. No single coefficient confirms a construct; each study adds to or weakens the case, and a new use or population reopens it. That makes construct validity easy to assert and slow to earn. The defensible claim names the concept, the evidence gathered so far and the gaps that remain, instead of calling the instrument validated and moving on.
In practice
A leadership team is offered a new measure of adaptability. Before comparing scores across candidates, the people team asks what the scores correlate with. The answer shows they track self-reported openness closely and observed behaviour under changing priorities barely at all, which suggests the scale captures how people describe themselves rather than how they adapt.
Evidence
Construct validity concerns how well a set of indicators reflects a concept that is not directly measurable.
Construct validity, Wikipedia (2026)Samuel Messick set out six distinguishable aspects of construct validity: content, substantive, structural, generalisability, external and consequential.
Samuel Messick, Validity of Psychological Assessment, Educational Testing Service Research Report RR-94-45 (1994)
What it cannot tell you
Construct validity evidence shows that scores reflect a concept in the populations and uses studied. It does not show that the scores predict a particular outcome, and it does not carry automatically to a new use, a new language or a new population, each of which needs evidence of its own.
Questions
Samuel Messick's 1994 report for Educational Testing Service names six: content, substantive, structural, generalisability, external and consequential. Together they ask whether the items cover the concept, whether the scores behave as the concept predicts, and what follows from using them. No single aspect settles the question alone.
Reliability asks whether a score is consistent across items, occasions or raters. Construct validity asks whether that consistent score reflects the intended concept. A measure can be highly reliable and still capture something else, such as reading speed or a wish to look good, so reliability is necessary but not sufficient.
By checking that scores correlate with measures of related concepts, correlate weakly with measures of unrelated ones, show the expected internal structure, and behave as theory predicts in groups known to differ. Each result adds to the case. None of them, on its own, establishes that the scores mean what their label says.
No. Evidence is gathered for an interpretation in a population. The 2014 Standards for Educational and Psychological Testing tie validity to the proposed uses of scores, so a new language, role level or decision reopens the question rather than inheriting the earlier answer from a different setting.
Because a decision about a person rests on what the score is taken to mean. If a scale labelled leadership mostly measures confidence, every appointment it informs carries that error, however consistent the scores are. Construct evidence is the check that the label on the scale matches what it actually captures.