Scored, feedback-driven forecasters improve; unscored experts performing the same task do not.
Why it matters when the plan changes
Professional judgement is almost never scored. Predictions are made, outcomes occur, and nobody puts the two side by side, so the same errors recur indefinitely with no mechanism to correct them. Forecasting only becomes a discipline that improves once predictions are recorded in advance and later set against what actually happened; without that step, scoring is unglamorous administration that turns out to be the entire difference between judgement that improves and judgement that recycles.
The tension is exposure. A scored forecast can be shown to have been wrong, publicly and specifically, which is why most advisory work avoids producing them. Proper scoring rules go further, measuring not just whether a call was right but how well the stated confidence matched the outcome, which is a harder thing to hide from. The willingness to be scored is the costly signal, and it is also the only way a supplier can eventually claim anything about accuracy with evidence behind it.
In practice
A judgement says an unowned interface will delay a milestone unless an owner is assigned within six weeks, and sets a review at eight. At the review the owner was assigned in week five and the milestone held. That outcome goes into the record, and the next similar call is made against something observed rather than remembered.
Evidence
Probabilistic forecasts are scored with proper scoring rules that measure the accuracy of stated probabilities.
Brier score, Wikipedia (2026)Forecasts recorded in advance can later be compared with actual outcomes, which is what scoring requires.
Forecasting, Wikipedia (2026)
What it cannot tell you
A scored forecast records accuracy for a specific claim and a specific horizon; it says nothing about calls outside that scope. A long run of correct scores does not confirm the reasoning behind them, only the outcome. Scoring measures calibration or hit rate; it cannot by itself tell you whether the underlying judgement generalises to a different domain.
Questions
Recording the prediction, its confidence and its review date before the outcome, then comparing the two when the review point arrives. Proper scoring rules such as the Brier score, described on Wikipedia in 2026, measure how well stated probabilities matched what happened, not merely whether the direction was right.
Because it requires committing to something specific in advance and then publishing when it was wrong. Descriptive advice avoids both. The absence of scoring is also why decades of experience in a field do not reliably produce better judgement in that field.
Enough that the pattern is not luck, which is tens rather than a handful. A gate of fifty scored forecasts before any accuracy claim, held to a number set in advance, is what stops the threshold moving once results start arriving.
It is recorded as wrong, alongside what the reasoning was and what the evidence had shown. That record is more useful than a correct one, because it is where the method changes. Quietly dropping incorrect calls is what turns a ledger into marketing.
It works better with them, because proper scoring rules measure how well stated confidence matched outcomes rather than only direction. Forecasting, as described on Wikipedia in 2026, notes that recorded predictions can later be compared with actual outcomes, which is what scoring requires even with coarser confidence levels.