Skip to content
Strategy Execution

The Unpriced Risk.

By Kevin Bjerring, Philip Myrup · 2026 · 11 min

Every strategic plan is priced for risk before it is approved. Financial risk is modelled across scenarios. Market risk is stress-tested against competitor response. Operational risk is decomposed into named dependencies with named owners and a mitigation for each. People risk appears as a single line, usually labelled change management, which is a mitigation rather than a measurement. This paper examines why the input that most determines whether a plan converts is the one input nobody prices. Personality measured without reference to a situation predicts job performance at a mean validity of .09. Measured against a specified demand, that figure more than doubles. People execution risk is not unmeasurable. It is unmeasured, and every instrument currently pointed at it is aimed at the wrong unit.

The Unpriced Risk.

A plan that reallocates capital arrives at the board with a sensitivity table. A plan that reallocates authority, sequencing and accountability across a leadership team arrives with a slide about communication.

The asymmetry is not a matter of sophistication. It is a matter of which risks have been decomposed. Financial risk has line items. Market risk has scenarios. People risk has a category name and nothing underneath it, so it is carried into execution as a single undifferentiated bet on the current team.

The aggregate consequence is documented and structurally uninformative. A Marakon Associates and Economist Intelligence Unit survey of senior executives at 197 large companies found that strategies delivered on average 63% of the financial performance they promised, with more than a third of respondents placing their own figure below 50% (Mankins & Steele, 2005). The authors' central observation is the one that matters. The processes companies use make it difficult to discern whether the shortfall came from poor planning, poor execution, both, or neither. The gap is quantified. Its composition is not.

Where the composition has been examined, it is people-shaped. Sull et al. (2015) surveyed 7,600 managers across 262 companies in 30 industries. Only 55% of middle managers could name even one of their company's top five strategic priorities. 84% reported they could rely on their boss all or most of the time, but only 9% said they could rely on colleagues in other units all of the time. Conflicts between units were handled badly two times out of three: resolved after significant delay 38% of the time, resolved quickly but poorly 14%, and left to fester 12%. Priority comprehension, cross-unit reliance and conflict resolution are people variables. None of them appears on a conventional risk register.

The absence has a diagnostic signature. Claims that 50% to 90% of strategic initiatives fail circulate widely, and a review of the published estimates concludes that the true rate remains undetermined, because the evidence behind those estimates is outdated, fragmentary, fragile or absent (Cândido & Santos, 2015). The reason no rate exists is that readiness is almost never measured before a programme and outcomes almost never measured against that baseline. A risk with no denominator cannot be priced, and a risk that cannot be priced is managed as though it were zero.

The Aggregate Instruments.

The absence of people risk pricing is not an oversight. It reflects the structural limitations of every instrument currently available. Each solves a version of the measurement problem. None solves the right version.

Engagement surveys measure sentiment about conditions that already exist. They do this well, and they are lagging by construction. Sentiment about a strategic change is a consequence of how that change is being led, which means the instrument registers the problem only after it has begun producing effects. Culture profiles describe the aggregate environment, which is useful for benchmarking and silent on how a specific executive experiences that environment. Talent reviews calibrate individuals against the organisational chart the new plan is about to redraw. Personality profiles usually already exist, taken at hire or at an offsite, debriefed across an afternoon and filed as descriptions of individuals.

That last use is the one the evidence identifies as least predictive. Tett and Burnett (2003) report that across all trait and criterion combinations, the mean corrected validity of personality for job performance is approximately .09, with a lower 90% credibility value of -.13. The same trait predicts performance positively in one setting and negatively in another. Measured without a situation, the instrument yields a description, not a forecast.

Measured against a situation, the validity returns. Shaffer and Postlethwaite (2012) meta-analysed 90 studies comparing general personality measures against work-contextualised ones. Noncontextualised measures produced validities from .02 to .22, with a mean of .11. Contextualised measures produced validities from .14 to .30, with a mean of .24. Sackett et al. (2022) treat contextualised and decontextualised personality measures as distinct predictors for the same reason. The frame of reference is not an administrative property of test delivery. It is the majority of the predictive power.

The timing compounds the error. Meyer et al. (2010) synthesise situational strength into four facets: clarity, consistency, constraints and consequences. Strong situations, where expectations are unambiguous and enforced, suppress the expression of individual differences. Weak situations allow disposition to drive behaviour. A strategy that has operated for five years is a strong situation. The first year of a new one is a weak situation by construction, because the old cues no longer apply and the new ones are not yet built. Pressure narrows attention to fewer cues (Easterbrook, 1959) and shifts behaviour from goal-directed toward habitual control, a mechanism with converging though not yet settled evidence in humans (Schwabe & Wolf, 2009). Individual patterns exert maximum influence at precisely the moment no instrument is pointed at them.

Capability is also a property of the person in a setting rather than a property carried between settings. Managers leave a measurable imprint on the firms they run (Bertrand & Schoar, 2003), yet firm policy changes following exogenous chief executive departures show no abnormal variability, which led Fee et al. (2013) to conclude that managerial style does not transfer across employers. Read together, the two findings say something more useful than either alone. Managers differ, and what they produce is not portable.

This is the limitation that constrained the dominant academic framework for four decades. Upper echelons theory predicts that organisations reflect their top teams (Hambrick & Mason, 1984), and it operationalised that prediction through executive demography as a measurement proxy for the cognitions it could not reach directly. Its own reviewers have asked the field to examine what demographic characteristics actually mean against the deeper constructs they are presumed to proxy (Carpenter et al., 2004). The profiles already sitting in most organisations measure those constructs. They have never been pointed at a plan.

The Capability-Energy Distinction.

People execution risk is not one risk. It resolves into two, and conflating them is what makes it look unmeasurable.

The first is capability against demand. Kaplan et al. (2012) assessed more than 300 chief executive candidates and found that execution-oriented characteristics predicted subsequent performance more strongly than interpersonal ones, and that success was at best marginally related to incumbency once observable ability was held constant. The dominant informal assessment process inverts this weighting, overweighting interpersonal impression and underweighting the dimension with more predictive power.

The second is sustained energy, and it is the one no assessment process reads. A leader may possess the capability to execute a pricing overhaul and derive no energy from that work. They will deprioritise it in favour of the pipeline they find interesting. The plan drifts. The model breaks. This happens not because the leader lacks skill but because nothing distinguished capacity from drive.

The fit literature confirms the split and locates it precisely. Kristof-Brown et al. (2005) synthesised 172 studies and found person-job fit strongly related to job satisfaction, at a corrected correlation of .56, and to intention to leave, at -.46, and only weakly related to job performance. Fit predicts attachment. It does not predict execution. Both readings matter to a plan and they answer different questions. One identifies whose pattern will collide with the new demand. The other identifies whose energy the plan will draw down, and therefore who will still be in the seat in eighteen months.

Two risk lines, not one. Where they diverge, execution risk compounds quietly and predictably.

The Composition Layer.

Individual readings, however rigorous, miss the layer where execution actually sits.

LePine (2003) studied 73 teams and found that members' individual differences explained no variance in team performance while the task remained routine. After an unforeseen change in the task context, teams whose members had higher cognitive ability, achievement orientation and openness, and lower dependability, adapted their role structure faster and performed better. Two findings there matter. The signal was invisible in steady state and appeared only when the plan changed. And lower dependability helped, which is the inverse of what a conventional scorecard rewards. The study is a simulation, which bounds the generalisation to executive teams, and it remains the cleanest available demonstration that the right composition for change is not the right composition for continuity.

At the executive level the pattern holds under real conditions. Virany et al. (1992) followed 59 firms through a turbulent industry and found that chief executive succession and executive-team change each independently improved subsequent performance, with the effect accentuated when either coincided with strategic reorientation. The rarer and more effective long-run pattern combined reorientation and team change while retaining the chief executive, so the team learned without discarding its institutional memory. That is a configuration judgement, and no individual reading supports it.

Configuration is also what travels. Groysberg et al. (2008) found that star analysts who changed firms suffered performance decline persisting at least five years, most acute among those who moved alone to weaker platforms. Those who moved with their teams, or to firms with stronger capabilities, showed no significant decline. Performance did not travel with the individual. It travelled with the system around them. The failure case is documented in matched pairs: Hambrick and D'Aveni (1992) compared 57 large bankruptcies with 57 surviving peers and found top-team composition diverging and accelerating over the final five years, with causality running in both directions.

The composition evidence carries one constraint an honest instrument must state. Deep-level composition, meaning members' personality, values and abilities rather than their demographics, does predict team performance. The effects are modest, they vary by setting, and the correct way to aggregate individuals into a team score differs by trait, with the team minimum mattering for some traits and the mean for others (Bell, 2007). Composition carries signal. How that signal is aggregated is a design decision with consequences, and it is the decision we are researching rather than one we regard as settled.

The practical failure is predictable. Four individually strong readings produce four green indicators. What the indicators do not reveal is that three of the four share the same strength profile and the same blind spot, that commercial capability is triple-covered, and that the workstream the plan depends on has no natural owner.

The Attribution Asymmetry.

If people risk is decomposable and the stakes are documented, the question becomes why the gap persists. It persists because the risk is priced retrospectively, and the retrospective price is always attached to a person.

Friction rarely announces its cause. A decision is slow, so the leader lacks pace. Perhaps. Or the authority for that decision is split between two roles and the evidence needed to make it arrives late. An executive is struggling, so the person is not capable. Perhaps. Or the role changed while the support around it did not, and success is still defined by the abandoned strategy. Every first explanation might be true. That is what makes them structurally dangerous. Each is flattering to the plan and cheap to act on, and each terminates the inquiry.

The substitute for measurement is experienced judgement, and the evidence does not support the substitution. Grove et al. (2000) meta-analysed 136 comparisons of clinical against mechanical prediction. Mechanical techniques were on average approximately 10% more accurate, substantially outperformed clinical judgement in 33% to 47% of studies, and were substantially outperformed in only 6% to 16%. The superiority held regardless of the judgement task, the type of judges, or the judges' amount of experience. Kahneman et al. (2021) document how professionals reach materially different conclusions from identical cases. Experience is not the problem. Unstructured experience is noisy, and noise is invisible from inside it.

Three mechanisms sustain the asymmetry. First, no organisation systematically tracks whether its pre-change reading of the leadership team predicted the post-change outcome, so the attribution is never falsified and the method never improves. Second, the cost of measurement is visible and certain while its value is diffuse and probabilistic. Third, the belief that management quality is already assessed through the normal process supplies a zero-cost alternative that avoids explicit expenditure.

The cost of the wrong attribution is not symmetrical. Removing a capable executive whose role was undeliverable retains the problem and consumes a year discovering that. The successor inherits the same role, the same split authority and the same measurement. Everyone watching learns what happens to whoever holds that seat, which changes who will accept it next and on what terms.

Conclusion.

The failure to price people execution risk is not a failure of awareness. Every organisation knows leadership determines whether a plan converts, and someone in the building can usually name the month the friction started. It is a failure of sequence. The risk is read after the outcome, when the only remaining variable is a person's continued employment, and the evidence that would have separated the person from the arrangement has been sitting in the organisation, unread against anything, since the day it was collected.

The instruments currently in use are not underdeveloped versions of the right approach. A reading of a leader's general disposition, absent a specified demand, predicts at .09. An engagement survey measures the friction after it has begun. A talent review calibrates against a chart the plan is about to redraw. Each is structurally incapable of answering the question the plan actually poses, which is what this specific strategy will demand of this specific role, what the person in it supplies, and where across the team those two things fail to meet.

The shift we are researching is from people risk described after the fact to people risk priced against a named plan: which patterns will collide with the new demand, whose energy the plan will draw down, and how those readings distribute across the team that has to carry it. The person is never the data point. A leader is not a score and a hard quarter is not a verdict. The discipline is to describe the arrangement around the person, with every source shown, and to leave the decision with the people accountable for it.

The question for any organisation approving a strategic plan is not whether it carries people risk. It is whether that risk was priced before the vote, or discovered in month nine.

References

Bell, S. T. (2007). Deep-level composition variables as predictors of team performance: A meta-analysis. Journal of Applied Psychology, 92(3), 595-615. https://doi.org/10.1037/0021-9010.92.3.595

Bertrand, M., & Schoar, A. (2003). Managing with style: The effect of managers on firm policies. The Quarterly Journal of Economics, 118(4), 1169-1208. https://doi.org/10.1162/003355303322552775

Cândido, C. J. F., & Santos, S. P. (2015). Strategy implementation: What is the failure rate? Journal of Management & Organization, 21(2), 237-262. https://doi.org/10.1017/jmo.2014.77

Carpenter, M. A., Geletkanycz, M. A., & Sanders, W. G. (2004). Upper echelons research revisited: Antecedents, elements, and consequences of top management team composition. Journal of Management, 30(6), 749-778. https://doi.org/10.1016/j.jm.2004.06.001

Easterbrook, J. A. (1959). The effect of emotion on cue utilization and the organization of behavior. Psychological Review, 66(3), 183-201. https://doi.org/10.1037/h0047707

Fee, C. E., Hadlock, C. J., & Pierce, J. R. (2013). Managers with and without style: Evidence using exogenous variation. The Review of Financial Studies, 26(3), 567-601. https://doi.org/10.1093/rfs/hhs131

Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., & Nelson, C. (2000). Clinical versus mechanical prediction: A meta-analysis. Psychological Assessment, 12(1), 19-30. https://doi.org/10.1037/1040-3590.12.1.19

Groysberg, B., Lee, L.-E., & Nanda, A. (2008). Can they take it with them? The portability of star knowledge workers' performance. Management Science, 54(7), 1213-1230. https://doi.org/10.1287/mnsc.1070.0809

Hambrick, D. C., & D'Aveni, R. A. (1992). Top team deterioration as part of the downward spiral of large corporate bankruptcies. Management Science, 38(10), 1445-1466. https://doi.org/10.1287/mnsc.38.10.1445

Hambrick, D. C., & Mason, P. A. (1984). Upper echelons: The organization as a reflection of its top managers. Academy of Management Review, 9(2), 193-206. https://doi.org/10.2307/258434

Kahneman, D., Sibony, O., & Sunstein, C. R. (2021). Noise: A flaw in human judgment. Little, Brown Spark.

Kaplan, S. N., Klebanov, M. M., & Sorensen, M. (2012). Which CEO characteristics and abilities matter? The Journal of Finance, 67(3), 973-1007. https://doi.org/10.1111/j.1540-6261.2012.01739.x

Kristof-Brown, A. L., Zimmerman, R. D., & Johnson, E. C. (2005). Consequences of individuals' fit at work: A meta-analysis of person-job, person-organization, person-group, and person-supervisor fit. Personnel Psychology, 58(2), 281-342. https://doi.org/10.1111/j.1744-6570.2005.00672.x

LePine, J. A. (2003). Team adaptation and postchange performance: Effects of team composition in terms of members' cognitive ability and personality. Journal of Applied Psychology, 88(1), 27-39. https://doi.org/10.1037/0021-9010.88.1.27

Mankins, M. C., & Steele, R. (2005). Turning great strategy into great performance. Harvard Business Review, 83(7/8), 65-72. https://hbr.org/2005/07/turning-great-strategy-into-great-performance

Meyer, R. D., Dalal, R. S., & Hermida, R. (2010). A review and synthesis of situational strength in the organizational sciences. Journal of Management, 36(1), 121-140. https://doi.org/10.1177/0149206309349309

Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040-2068. https://doi.org/10.1037/apl0000994

Schwabe, L., & Wolf, O. T. (2009). Stress prompts habit behavior in humans. Journal of Neuroscience, 29(22), 7191-7198. https://doi.org/10.1523/JNEUROSCI.0979-09.2009

Shaffer, J. A., & Postlethwaite, B. E. (2012). A matter of context: A meta-analytic investigation of the relative validity of contextualized and noncontextualized personality measures. Personnel Psychology, 65(3), 445-494. https://doi.org/10.1111/j.1744-6570.2012.01250.x

Sull, D., Homkes, R., & Sull, C. (2015). Why strategy execution unravels, and what to do about it. Harvard Business Review, 93(3), 58-66. https://hbr.org/2015/03/why-strategy-execution-unravelsand-what-to-do-about-it

Tett, R. P., & Burnett, D. D. (2003). A personality trait-based interactionist model of job performance. Journal of Applied Psychology, 88(3), 500-517. https://doi.org/10.1037/0021-9010.88.3.500

Virany, B., Tushman, M. L., & Romanelli, E. (1992). Executive succession and organization outcomes in turbulent environments: An organization learning approach. Organization Science, 3(1), 72-91. https://doi.org/10.1287/orsc.3.1.72