Skip to content

Item response theory

Item response theory is the modern framework for designing, analysing and scoring tests. It models the relationship between a person's position on an underlying trait and how they respond to each individual item, which allows scoring that accounts for the difficulty and discrimination of the items themselves.

It is a paradigm for the design, analysis and scoring of tests, which is what separates it from simply adding up correct answers.

Why it matters when the plan changes

Classical scoring treats every item as equally informative, which they are not. Some items separate people well and some barely at all, and a score that ignores the difference wastes information. Item response theory is a paradigm for the design, analysis and scoring of tests because it models each item's contribution directly. This lets an instrument report how precisely it has measured a particular person, rather than one reliability figure for everyone, since reliability is the freedom of a measure from measurement error and can be estimated per person rather than once for a whole scale.

The tension is between rigour and legibility. The modelling is what makes modern instruments defensible, including forced-choice formats where Thurstonian item response theory recovers normative scores from rankings that resist faking, and it also makes them harder to explain to the people they are used on. An assessment nobody can explain will be distrusted whatever its properties, which is a real cost rather than a presentational one.

In practice

Two respondents get the same number of items in the same direction on a scale. One endorsed items that most people endorse; the other endorsed items that few do. Classical scoring reports them as equal. Item-level modelling places them in different positions, because the items themselves carried different amounts of information.

Evidence

What it cannot tell you

Item response theory models the relationship between trait and response given the items in a particular instrument; it cannot tell whether those items measure a trait that matters for a given role or outcome. Fit statistics judge how well a model describes the data collected, not whether the underlying construct predicts performance, which depends on separate evidence such as person by context research.

Questions

Classical test theory reports one reliability figure for everyone, as reliability is defined generally as freedom from measurement error (Reliability (statistics), Wikipedia, 2026). Item response theory models each item separately, so scoring accounts for difficulty and discrimination, and precision can be reported per person rather than once for the whole test.

Estimating how the probability of a particular response changes as a person's trait level changes. Item response theory is described, in the Item response theory entry on Wikipedia (2026), as a paradigm for the design, analysis and scoring of tests, precisely because each item's contribution is modelled rather than averaged away.

Because raw forced-choice rankings are within-person and cannot be compared between people. The Thurstonian variant applies item response modelling to the comparisons, recovering scores on a common scale. Without it the format would resist faking and produce results nobody could use for selection.

More than classical approaches, yes, which is a real constraint for a new instrument. That is one reason a norm base is built deliberately over time rather than assembled from whoever is available, and why early-stage instruments should be explicit about the sample their scoring rests on.

Indirectly but genuinely. It determines whether scores can be compared between candidates and how precisely each person has been measured. A buyer does not need the mathematics, and should ask which model is used and what sample it was estimated on.