A norm-referenced test estimates a person's position within a predefined population, so the population defines what the score means.
Why it matters when the plan changes
The same raw result becomes a different conclusion depending on the comparison group, because a norm-referenced test only ever yields an estimate of position within whichever population was chosen as the reference. Being in the top decile of the general working population and being in the top decile of senior leaders are very different statements, and a report frequently shows one while the reader assumes the other. The group is often the least prominent information on the page.
The tension is between relevance and size. The most relevant comparison group for a particular senior role may contain very few people, which makes the norms unstable. A large general population gives stable norms and a less meaningful comparison. This mirrors why predictive validity evidence, established for one use and one population, does not transfer to another; a norm group carries the same restriction. Every instrument makes this trade, and honest ones state where they landed.
In practice
A candidate for a chief executive role is reported at the eightieth percentile on a scale. The comparison group is the general working population. Against other people who have held the role, they might sit at the median. Nothing in the report is false, and the reader's inference is wrong.
Evidence
A norm-referenced test yields an estimate of the position of the tested individual within a predefined population on the trait being measured.
Norm-referenced test, Wikipedia (2026)Validity evidence is established for a specific use and population and does not transfer to another.
Predictive validity, Wikipedia (2026)
What it cannot tell you
A norm group tells you where a score sits relative to a chosen population; it says nothing about whether that population is the right one for the decision, and nothing about absolute skill or suitability. Two people with identical raw performance can receive opposite comparative scores depending only on which group each was measured against.
Questions
Who is in it, how many, when it was gathered, and whether it matches the population the norm-referenced test defines, as described in Norm-referenced test, Wikipedia (2026). A percentile reported without that information cannot be interpreted, though it often is anyway.
Large enough for stable estimates, which depends on the model and how finely the group is subdivided. Splitting by role, level and country multiplies the requirement quickly, which is why a headline sample size says less than the size of the specific subgroup a person is compared against.
Yes. Populations change, roles change and the people who complete an instrument change as it spreads. Norms gathered years ago describe a different working population, and a supplier should say when theirs were collected and how often they are refreshed.
Against the group that makes the decision meaningful, which is usually the more specialised one, provided it is large enough to be stable. Where it is not, the honest approach is to report against the broader group and say plainly that the comparison is broad.
Closely. Validity evidence, as defined in Predictive validity, Wikipedia (2026), is established for a specific use and population and does not transfer to another; the same is true of norm groups. If a norm group under-represents a population, scores are positioned against a comparison that does not describe them.