Without this modelling step, forced-choice scores describe a person against themselves and nothing else.
Why it matters when the plan changes
Forced choice resists distortion and produces the wrong kind of number. Ordinary self-report is vulnerable to faking and response style; ranking statements against each other closes that route, but a ranking only records preference within one respondent. Rankings cannot support a comparison between two people without a model that reconstructs the dimensions underneath the comparisons, extending the older logic of scaling through repeated pairwise comparisons rather than absolute ratings. That reconstruction is the entire reason the format became usable in high-stakes settings.
The tension is between defensibility and legibility. The modelling, set within item response theory's framework for the design, analysis and scoring of tests, is what makes the scores comparable, and it also makes the instrument harder to explain to the people it is used on. An assessment that cannot be explained plainly will be distrusted regardless of how well it is constructed, which is a real cost rather than a presentational one.
In practice
Two candidates both rank decisiveness above consultation. Taken as rankings, the two results are identical. Modelled, they are not: one ranks decisiveness marginally higher across a narrow range and the other decisively, across a wide one. The comparison the decision needs exists only after the modelling step.
Evidence
The approach builds on scaling through repeated pairwise comparisons rather than absolute ratings.
Law of comparative judgment, Wikipedia (2026)Item response theory is the framework used for the design, analysis and scoring of such instruments.
Item response theory, Wikipedia (2026)
What it cannot tell you
Thurstonian IRT recovers comparable scores from rankings, but it cannot make the underlying statements equally desirable to begin with, nor can it substitute for validity evidence about what the scores predict. It says nothing about whether the dimensions modelled are the right ones for a given decision, only that the comparisons are usable across people.
Questions
The modelling treats each forced choice as a comparison between two statements, extending the logic of the Law of comparative judgment, which scales stimuli through repeated pairwise comparisons (Wikipedia, 2026). It works backwards to the dimensions explaining the pattern of choices, producing scores that can be compared between people, which raw rankings cannot support.
Item response theory, described in 2026 sources as a paradigm for the design, analysis and scoring of tests, exists precisely because a ranking says only which option a person preferred, not how strongly or against what baseline. Two people with identical rankings can differ substantially on the underlying dimensions.
It makes the result usable for a comparison it otherwise could not support. Accuracy is a separate question answered by reliability, validity and fairness evidence for the specific use, which is a higher bar and one that has to be met with its own evidence.
It builds on the work of L. L. Thurstone on comparative judgement in the 1920s, later combined with modern test theory to model forced-choice questionnaires. The Atlas instrument was built with Dr Susanne Frick at TU Dortmund, who researches this area and forced-choice design.
It is in validation. Reliability thresholds were pre-registered per construct before the pilot, so the success bar cannot be adjusted after results exist, and validity and fairness work is planned for the specific intended use. No claim of peer-reviewed or academic validation is made until one exists.