Skip to content

Reliability

Reliability is the extent to which a measure is free from measurement error, so that it gives consistent results when the thing being measured has not changed. It is checked across occasions, across items within a scale, and sometimes across raters. Without it, no other claim about an instrument can be assessed.

Reliability is the freedom of a measure from measurement error, which makes it the precondition for validity rather than a competing virtue.

Why it matters when the plan changes

An unreliable score changes when the person has not. Any decision resting on it rests partly on noise, and noise is invisible because a single number looks the same whether it is precise or not. In classical test theory, reliability is the share of observed variation that reflects true variation rather than error, which is why it is checked across occasions, items and raters. Reliability is the first question to ask of an instrument, and the one most often skipped, because the answer is a table, not a story.

The tension is that reliability is necessary and not sufficient. A perfectly consistent measure of something irrelevant is still useless, and a measure can be made more reliable simply by narrowing it until it no longer captures what matters. Validity evidence, the case that a score predicts what it claims to, is established for a specific use, and reliability is what makes that evidence interpretable. High reliability is a floor to clear before asking what the score predicts.

In practice

A candidate retakes an assessment eight weeks apart and scores differently enough to change the conclusion drawn about them. Nothing about the candidate changed in that period. The instrument's test-retest figures were never published, and the decision made on the first result rested on a number that would not repeat.

Evidence

What it cannot tell you

Reliability tells you a score is consistent, not that it measures anything useful. A perfectly stable instrument can still measure an irrelevant construct, and reliability alone says nothing about whether the score predicts performance, behaviour or outcome. It is a precondition for validity, not a substitute for it.

Questions

Test-retest, whether the same person scores similarly on two occasions. Internal consistency, whether items in a scale agree with each other. Inter-rater, whether different raters reach the same judgement. As Reliability (statistics), Wikipedia (2026) puts it, reliability is the freedom of a measure from measurement error, and which type is reported should match how the instrument is used.

It depends on the stakes and the construct. Higher bars apply where a score informs a decision about an individual than where it informs a group-level view. What matters more than any single figure is that the threshold was set before the data existed rather than chosen to fit the result.

Because a bar set after results exist is not a bar. Pre-registration fixes what would count as success before anyone knows the answer, which is the difference between testing an instrument and describing it favourably. It also makes a failed construct visible rather than quietly rescoped.

Yes, and this is the common trap. A measure can be made highly consistent by narrowing what it captures until it reliably measures something that does not matter. Predictive validity, per Predictive validity, Wikipedia (2026), is the extent to which a score predicts an outcome, and reliability alone says nothing about that.

Reliability figures per scale rather than an overall claim, the sample they were computed on, the interval used for test-retest, and whether the thresholds were set in advance. A supplier who answers with a single adjective rather than a table has not done the work.