Why the Best Decisions Go Uncounted
Ask a finance team to show you the cost of a bad decision and they can, eventually, with enough forensic work: the failed launch, the six-figure mis-hire, the platform rebuild. Ask them to show you the value of a good one and they go quiet, because a good decision leaves no trace. The system that did not break, the customer who did not churn, the regulation you were already ready for - none of it lands on a dashboard. The return on deciding well is real, large, and almost entirely invisible, which is precisely why it is the last thing most companies invest in improving.
The research that does exist is blunt about the size of the prize. Bain, in a 10-year research program covering more than 1,000 companies, found a clear correlation (at a minimum 95% confidence level) between how well an organization makes and executes decisions and how well it performs; in the survey behind its book Decide & Deliver, more than 750 companies, the relationship held in every industry and country studied. McKinsey put a price on the other side of the ledger. In a 2018 survey of about 1,250 respondents only 20% said their organization excels at decision making, and McKinsey modelled the cost of inefficient decision-making at a typical Fortune 500 company at around 530,000 days of managers' time a year, roughly $250 million in wages. That figure is a model built on stated assumptions, not a measured loss. The same survey found that respondents at organizations that decide both well and fast were twice as likely to report financial returns of 20% or more from their most recent big decision.
So the contribution of good decisions is not a soft, unmeasurable thing. Bain's data links it clearly to financial performance, though correlation does not show which way the effect runs. The problem is purely one of instrumentation: nobody is counting. This essay is the instrument.
The Measurement Principle: Two Layers, One Compounding Curve
You cannot measure a decision by asking whether it felt right. You measure it in two layers, separated in time. The leading layer is the quality of the choice while it is still cheap to change - is this the right feature, the right hire, the right sequence, the right structure. The lagging layer is the money that quality turns into later - development cost avoided, turnover cost avoided, time-to-market gained, defects that never reached a customer. Leading signals are early and directional; lagging metrics are late and financial. A measurement model that only watches the lagging layer is just an autopsy. The value is in reading the leading layer early enough to act.
Decision quality
Right feature, right hire, right sequence - measured while it is still cheap to change.
Operational result
Adoption, retention, cycle time, defect escape - the outcome the choice produced.
Dividend or debt
Cost avoided or cost incurred, priced in your own unit economics.
The second half of the principle is that the gap between a good and a bad decision does not stay the same size - it compounds. A structural or planning gap that would cost a modest amount to fix at thirty people costs far more once the same company is at a few hundred, and more again at a thousand, because everything built on top of the gap has to be unwound. This is the same curve software engineering has measured for decades: a defect caught at requirements is cheap; the same defect caught in production costs, in an illustrative table in NIST's 2002 report, up to 30 times as much to fix once released (10 times at system test). Decisions have an identical cost-of-change curve. Measuring the leading layer is how you catch the gap on the cheap side of it. (The organizational version of this curve is its own essay - see The 10x Rule.)
What follows are three ledgers where decision quality compounds into money. Each one gives you the leading signal to watch, the lagging metric it turns into, a benchmark to compare against, and a formula to price it in your own numbers - because a real measurement is your data, not a borrowed statistic.
Ledger One - Build: The Right Things, Well, and Changeably
Most build measurement stops at velocity - how much did we ship. Velocity is the wrong question, because the most expensive work a team does is the work that should never have been built. Pendo, analysing 615 software subscriptions, found that roughly 80% of features in the average product are rarely or never used, and that an average of 12% of features generate 80% of daily usage. Every one of those unused features was a decision - a planning choice about what mattered - and it was paid for in full: designed, built, tested, maintained, and carried as complexity forever. The build ledger measures whether you are choosing the right things, building them well enough to not pay twice, and keeping the product changeable when circumstances move.
Build
Are we making the right things, well, and in a way we can change later?
| Leading signal | Lagging money metric | Benchmark to compare |
|---|---|---|
| Feature adoption rate | Engineering spend on features that miss adoption | About 12% of features generate 80% of usage (Pendo) |
| Rework ratio | Cost of redoing vs. building new | Track your own trend; rising rework is the tell |
| Change failure rate & lead time | Cost of failed changes and slow delivery | Elite performers: about 5% change failure rate (DORA 2023) |
| Defect escape rate | Late-defect premium (up to 30x, NIST example) | NIST / Boehm cost-of-change curve |
| Iterations-to-value | Cycles burned before a feature works | Fewer iterations = tighter upstream decisions |
Two of these deserve emphasis because they are the ones that decide whether the company can survive its own success. Lead time for changes and change failure rate - two of the four DORA metrics that a decade of delivery research has linked to software delivery performance - together measure whether you can move fast without breaking things. And iterations-to-value, how many cycles a team burns before a feature actually works, is the cleanest proxy there is for upstream decision quality: tight decisions ship in one or two passes; vague ones churn for five. The dividend here is not abstract. At the digital bank where I worked, the onboarding I built carried an R&D team from 30 to 150 developers.
Ledger Two - People: Retention, Ramp, and Knowledge
People decisions are where invisible value hides most expensively, because the cost of a bad one is both large and delayed. Gallup estimates that replacing an employee costs between one-half and two times their annual salary, and more for managers and leaders. SHRM's 2025 benchmarking puts the median direct cost-per-hire at $1,200 for nonexecutive roles and about $10,600 for executives, with a median time-to-fill of about a month and a half; an earlier SHRM benchmark (2022) put the average at nearly $4,700. Those are direct costs only, and SHRM also cites employer estimates that the total cost of a new hire can reach three to four times the position's salary. Every one of those numbers traces back to a planning decision that was made, or skipped, months earlier.
People
Are we keeping the people we want, ramping them fast, and not losing the knowledge when they move?
| Leading signal | Lagging money metric | Benchmark to compare |
|---|---|---|
| Regretted attrition rate | Turnover & rehiring cost | 0.5x-2x annual salary per exit (Gallup) |
| Time-to-productivity | Ramp cost of every new hire | Strong onboarding lifts productivity 70%+ (Brandon Hall) |
| 90-day new-hire retention | Cost of early, avoidable exits | Strong onboarding lifts retention 82% (Brandon Hall) |
| Knowledge concentration (bus factor) | Exposure when one person leaves | Any critical process with a single owner |
| Cross-department handoff loss | Rework from knowledge that did not transfer | Track handoffs that get redone downstream |
The two under-measured signals here are time-to-productivity and knowledge concentration. Time-to-productivity is where the onboarding decision cashes out: Brandon Hall Group's research, as widely reported, found that a strong onboarding process improves new-hire productivity by over 70% and retention by 82% - meaning the choice to build a real onboarding path, rather than leave people to absorb the job by osmosis, is one of the highest-return planning decisions a growing company can make. Knowledge concentration - the bus factor - measures how much of the company lives in one person's head and walks out the door with them. It rarely appears on any report until the day it becomes a crisis.
Ledger Three - Alignment: When Functions Row Together
The third ledger is the one no single department owns, which is exactly why it leaks the most. Product decides one thing, marketing promises another, customer service is briefed on a third, and risk and legal find out when it is already shipped. Each function can be locally excellent and the company can still be slow, because the cost lives in the seams between them - the launch marketing was not ready for, the feature support could not explain, the compliance issue caught after go-live instead of before. This is where McKinsey's modelled $250M of wasted decision time would actually accumulate: not in any one function, but in the handoffs and the re-decisions between them.
Alignment
Do product, marketing, customer service, risk, and legal move as one system, or re-litigate every decision?
| Leading signal | Lagging money metric | Benchmark to compare |
|---|---|---|
| Decision cycle time | Manager-hours lost to slow / repeated decisions | ~530K days & $250M/yr at a Fortune 500 (McKinsey model) |
| Cross-functional rework | Work redone because functions were not aligned | Track launches reworked after release |
| Launch readiness | Cost of launches sales/CS could not support | Marketing & CS briefed before, not after |
| Risk & legal lead time | Cost of issues caught late vs. designed in | Same late-fix premium applies (up to 30x, NIST example) |
| Overall decision effectiveness | The financial correlate of all of the above | Clear correlation with performance, 95% confidence (Bain) |
The measurable heart of this ledger is decision cycle time: how long it takes, on average, to get a real cross-functional decision made and to make it stick. McKinsey's finding that speed and quality rise together - respondents whose organizations are good at both were twice as likely to report 20%+ returns on a big decision - means cycle time is not a trade-off against quality; it is a symptom of the same underlying alignment. When product, marketing, customer service, risk, and legal share one scoreboard and clear decision rights, decisions get faster and better at once. When they do not, every choice is re-opened, and the re-opening is the cost.
The Cheapest Signal Is the One Culture Hides
Every signal in the three ledgers assumes one thing: that the truth actually reaches you. Culture is what decides whether it does, which makes it the cross-cutting ledger - it can corrupt all fifteen signals at once, and it is usually the cheapest to fix, because the repair is a changed behaviour, not a new system. It is also the part you cannot instrument from a dashboard. You read it by observing, and the return on reading it well is wildly out of proportion to what the fix costs.
Start with what happens when a culture punishes the person who raises the problem. The problem does not go away; it simply stops being raised. Google's Project Aristotle, studying 180 teams, found psychological safety - the sense that you can speak up without being punished - to be the most important of the five dynamics that set its most effective teams apart. Research on employee silence shows the shadow side of the same thing: in an interview study of 40 employees (Milliken, Morrison and Hewlin, 2003), 85% said they had been in a situation where they felt unable to raise an important issue with a supervisor. It is a small sample, and it points the same way. That silence has a visible shape - status reports where everything is green, meetings where people argue for the answer the manager already prefers, bad news that arrives late or pre-softened. It is the exact mechanism that makes a bad decision invisible until it is expensive: the leading signal existed, and the culture filtered it out before anyone could act on it while it was still cheap.
The same culture quietly exports your best judgement. The people who see the problems most clearly are often the least willing to stay quiet about them, so they are the first to leave. MIT Sloan's analysis of the 2021 attrition wave found a toxic culture to be about 10 times more powerful than compensation in predicting who quits; Gallup found that managers account for at least 70% of the variance in engagement scores across business units, and that one in two U.S. adults it surveyed had left a job to get away from a manager. You lose the judgement before you lose the headcount, and the exit interview almost never names the real cause.
The founder version of this sits at the top of the chart. When every decision of consequence routes through one person, usually the CEO or founder, the company can look fast and coherent - decisions are quick, the vision is consistent, nothing appears to need fixing. The fragility is invisible precisely because that person is still in the room. It surfaces the moment they try to step back to work on strategy, or take a real break, and the delegation, the decision rights, and the documented judgement turn out never to have been built. The business was running on one person, not one system - a single point of failure identical to the bus factor on the people ledger, just harder to say out loud. Building real delegation before that moment, rather than during the crisis it causes, is one of the highest-return structural decisions a founder-led company ever makes.
None of this needs a platform or a budget to find, and that is the point. It shows up in an afternoon of watching how the organization actually behaves: who speaks and who goes quiet, which reports are suspiciously smooth, how many decisions stall until one calendar frees up, how people talk about a mistake. The fixes are usually small - an explicit decision right that lets someone act without asking, a meeting rule that rewards whoever surfaces a risk, a standing decision the founder deliberately stops making - and the payoff is large, because you are repairing the channel every other measurement depends on. This is the highest-leverage, lowest-cost work there is, which is exactly why a real assessment is built on observation, not just dashboards: a dashboard reports the numbers a culture is comfortable showing, and the expensive ones are usually the ones only observation catches.
The Planning Health Scorecard
Three ledgers, fifteen signals. To make them usable you compress them into a single one-page instrument you can score in an afternoon and re-score every quarter. For each ledger, rate the leading signals from 1 (no visibility, no discipline) to 5 (measured, benchmarked, acted on), attach the one lagging money figure you can defend, and write down the single most expensive gap. The score tells you where the decision debt is concentrated; the money figure tells you what closing it is worth; the trend, quarter over quarter, tells you whether your decisions are actually getting better.
| Ledger | Score the signal (1-5) | Attach the money | The question it answers |
|---|---|---|---|
| Build | Adoption, rework, change failure, defect escape, iterations-to-value | Wrong-feature + late-defect cost per year | Are we spending build money on things that matter and last? |
| People | Regretted attrition, time-to-productivity, 90-day retention, bus factor | Turnover + onboarding drag per year | Are we keeping and ramping the people we chose? |
| Alignment | Decision cycle time, cross-functional rework, launch readiness, risk lead time | Decision drag + rework cost per year | Do the functions decide once, together, and move? |
The output is deliberately not a single vanity index. A blended "planning health = 3.8" number hides exactly the thing you need to see, which is which ledger is bleeding. The value of the scorecard is the ranking it produces: three ledgers, sorted by the money attached to the gap, so the most expensive decision debt is the first thing you pay down. That ranking is the entire point of measuring - not to admire the score, but to know what to fix first.
A 30-Day Way to Run It Yourself
None of this requires a consulting engagement to start, and it should not. The default and the better outcome is a company that can read its own decision quality with the data and people it already has. Here is the version you can run in a month.
Week 1 - Baseline what you already have. You are further along than you think. Adoption data lives in your product analytics, change failure and lead time in your deployment tooling, attrition and time-to-fill in HR, decision friction in the calendars of your managers. Pull the fifteen signals as they stand, even if half are rough. The gaps you cannot measure yet are themselves a finding.
Week 2 - Price the top gap in each ledger. Take the single worst signal per ledger and run its formula on your real unit economics - your engineering cost, your salaries, your manager hours. You are not after precision; you are after an order of magnitude you can defend. "This is costing us somewhere around a quarter of a million a year" changes the conversation even at plus-or-minus 30%.
Week 3 - Sequence by money, not by noise. Rank the three priced gaps. The loudest problem is rarely the most expensive one, and the most expensive one is where the next decision should go. Wherever possible, the fix should be something your existing team can execute with what it already has - a changed process, a clearer decision right, a real onboarding path - not a new vendor, a new hire, or a tool you will spend a year integrating.
Week 4 - Set the re-score date. Put the scorecard on a quarterly cadence and assign one owner. A single measurement is a snapshot; the trend is the actual instrument. Decisions get better when someone watches whether they are getting better.
Where This Becomes the Assessment
Run the four weeks and you will have a defensible read on where your decisions are leaking value. Most companies get most of the way there on their own, and that is the intended outcome - the point of a measurement model is to leave you able to use it. Where I come in is when the leaks cross functions, when the numbers need pressure-testing against how comparable companies actually behave, or when reading the organization honestly is easier for someone who does not sit inside its politics.
That is exactly what the Scale Readiness Assessment is - this measurement run for you, end to end. It finds where decisions are quietly leaking value across the three ledgers, prices each leak in your own numbers, and, critically, matches the fix to your resources, market, and identity rather than handing you a generic playbook. A regulated fintech, a fast-moving startup, and a five-hundred-person operation all score the same three ledgers, but the right move for each is different - because the constraint, the risk tolerance, and the culture are different. Finding the leak is the seeing; matching the fix to what the company can actually run is the part that makes it get implemented. The engagement hands you a ranked, costed list of what to fix and what to prevent - built, wherever possible, to run on the team and budget you already have, with no redundant retainer attached.
This is the core of how I work: see the leak before it compounds, manage it down with the lightest change that holds, and build the structure so it does not come back. A good decision will never announce itself on your dashboard. The next best thing is a way to measure its absence before it gets expensive - which is what this model, and the assessment behind it, are for.
Where the Numbers Mislead
A measurement model is a tool, and tools cut both ways. The most common failure is optimizing the metric instead of the decision - a team that is measured on feature adoption will ship safe, incremental features and stop taking the bets that adoption cannot predict; a team measured on decision speed will make fast, shallow calls. Every signal here is a proxy, and a proxy chased hard enough stops being one.
The second caution is the benchmarks. The figures in this essay - Pendo's 80%, DORA's failure-rate bands, Gallup's 0.5x-to-2x, NIST's illustrative 30x - come from published sources, but they are cross-industry ranges, not your number. A regulated business should spend more on defect prevention and slower, more controlled decisions than a consumer app; measured against a startup's benchmark it will look wasteful when it is being correctly careful. Use published figures to frame the question and size the order of magnitude. Use your own data to answer it. The moment a borrowed statistic becomes the headline instead of your own priced gap, the measurement has stopped being honest - and an honest read of your own numbers is the only version worth acting on.
Sources
- Marcia Blenko, Michael Mankins, Paul Rogers, Decide & Deliver (Bain & Company) - data from more than 750 companies in the book; Bain's 10-year program of more than 1,000 companies reports a clear correlation, at a minimum 95% confidence level, between decision effectiveness and business performance. The 95% is a confidence level, not the strength of the correlation. bain.com, bain.com
- McKinsey & Company, Decision Making in the Age of Urgency (2019, survey fielded February 2018, about 1,250 respondents) - only 20% say their organization excels at decision making; ~530,000 days of managers' time (~$250M in wages) a year for a typical Fortune 500 is a modelled estimate; respondents at organizations that decide well and fast were twice as likely to report 20%+ returns on their latest big decision. mckinsey.com
- Gallup, This Fixable Problem Costs U.S. Businesses $1 Trillion (2019) - voluntary turnover costs U.S. businesses about $1 trillion a year; replacing an employee costs one-half to two times annual salary. gallup.com
- Nicole Forsgren, Jez Humble, Gene Kim, Accelerate / DORA - the four keys (deployment frequency, lead time for changes, change failure rate, time to restore) as measures of delivery performance. dora.dev
- DORA (Google Cloud), Accelerate State of DevOps Report 2023 - change failure rate by cluster: elite 5%, high 10%, medium 15%, low 64%. dora.dev
- Pendo, 2019 Feature Adoption Report - across 615 Pendo subscriptions, 80% of features in the average product are rarely or never used; an average of 12% of features generate 80% of daily usage volume. pendo.io
- NIST, The Economic Impacts of Inadequate Infrastructure for Software Testing (2002) - Table 5-1, an illustrative example: 1x at requirements, 5x at coding, 10x at integration and system test, 15x at beta, 30x after release. It builds on earlier Boehm and Baziuk studies; treat it as directional. nist.gov
- Society for Human Resource Management (SHRM), 2025 Recruiting Benchmarking Report - median cost-per-hire $1,200 (nonexecutive) and about $10,600 (executive); median time-to-fill about a month and a half; direct costs only. shrm.org
- SHRM, The Real Costs of Recruitment (April 2022) - average cost per hire of nearly $4,700; many employers estimate the total cost of a new hire at three to four times the position's salary. shrm.org
- Brandon Hall Group - as widely reported, a strong onboarding process improves new-hire retention by 82% and productivity by over 70%. brandonhall.com
- Donald Sull, Charles Sull, Ben Zweig, Toxic Culture Is Driving the Great Resignation (MIT Sloan Management Review, 2022) - toxic culture is 10.4 times more powerful than compensation in predicting attrition (1.4 million Glassdoor reviews). sloanreview.mit.edu
- Google, Project Aristotle (re:Work) - across 180 teams, psychological safety was the most important of five dynamics that distinguished the most effective teams. rework.withgoogle.com
- Frances Milliken, Elizabeth Morrison, Patricia Hewlin, An Exploratory Study of Employee Silence, Journal of Management Studies 40(6), 2003 - interviews with 40 employees; 85% said they had felt unable to raise an important issue with a supervisor (small qualitative sample). wiley.com
- Gallup, State of the American Manager (2015) - managers account for at least 70% of the variance in engagement scores across business units; one in two U.S. adults surveyed had left a job to get away from a manager. gallup.com
Related Reading
- The 10x Rule - the compounding cost-of-change curve behind this model: why a gap caught early costs a fraction of the same gap caught after a growth event.
- The Assessment Is the Product - what the Scale Readiness Assessment actually delivers, and why the diagnosis is the deliverable.
- Ship KPIs Like Features - how to build the scoreboard that makes decision quality visible in the first place.
- The KPI Trap - why the wrong metric on the scorecard quietly pulls a company in the wrong direction, and how to choose ones that hold.
- The GEAR Model - the operating structure that gives cross-functional decisions clear owners and one shared scoreboard.
May Mor
Interim and Fractional Product and Program Manager. M.Sc in AI, former Technical PM at a digital bank, where I built the onboarding that carried an R&D team from 30 to 150 developers, and at an adtech company. I find where decisions are quietly leaking value, price the leak, and match the fix to what your team can actually run. Full bio →