A record that contains no incorrect calls is a case-study list rather than a track record.
Why it matters when the plan changes
Every advisory supplier has references and case studies; almost none has a record that could show them to have been wrong. That asymmetry is why buyers cannot distinguish between firms on evidence and fall back on brand, relationships and price. A scored record, of the kind produced when predictions are checked against outcomes using something like a Brier score, is the only thing that breaks the tie with information rather than reputation.
The tension is elapsed time. A track record cannot be bought, accelerated or assembled from prior work; it accumulates at the rate that outcomes occur, which for organisational judgements is quarters, since a forecast can only be compared with an outcome once that outcome has actually happened. Any supplier claiming one early is describing something else, and any supplier building one honestly will be unable to show much for a while.
In practice
A buyer asks two suppliers for evidence of accuracy. One supplies eight case studies in which the work went well. The other supplies the number of judgements made, how many have reached their review point, and how those resolved, including the ones that did not hold. Only the second answered the question.
Evidence
Scored, feedback-driven forecasters improve, and the scoring is what produces the record.
Superforecasting and the Good Judgment Project (2026)Forecasts can be compared with actual outcomes, which is the comparison a track record preserves.
Forecasting, Wikipedia (2026)
What it cannot tell you
A track record shows what a forecaster has been right and wrong about within the classes of judgement it covers. It says little about performance on a materially different class of decision, and a short record cannot distinguish genuine skill from a run of favourable circumstance.
Questions
Case studies are selected after the outcome is known. A track record includes every judgement made, recorded before the outcome, with the ones that failed still in it, in the same spirit as the practice described in Superforecasting and the Good Judgment Project, where predictions are scored rather than curated afterwards.
As long as the outcomes themselves take to occur, which for organisational judgements is several quarters each. This mirrors the basic definition used in Forecasting, Wikipedia (2026), where a forecast can only be compared with an outcome once that outcome has actually happened, not before.
How many judgements have been made, how many have reached their review point, how those resolved, and whether the incorrect ones are in the record. A supplier who cannot answer the last question does not have a track record in the sense that matters.
Partially. Accuracy on one class of judgement says something about method and discipline and less about performance on a different class. A record should state what kinds of call it covers, in the same way validity evidence states its intended use.
Because it exposes the supplier to being demonstrably wrong, and because the business model of most advisory work does not require it. The absence is structural rather than accidental, which is also why building one is a durable position rather than a feature to copy.