Decision reliability for product & strategy

Every decision either
speeds you up
or slows you down.

Launch the product in March or in September. Give the project another quarter or stop it today. Roll the new way of working out to every squad or to one. The hard part is telling which is which, while you still have the choice — and your success is the sum of how those calls land. Reason Based Product Management gives everyone in the room one way to work that out: transparent while you decide, open to revision while you deliver.

Δ COMPANY VALUE DECISION IF YOU DO IT IF YOU DON’T
Fig. 01 — The same quarter, played two ways. The distance between them is what the decision is worth.
How success accumulates

Success is a few hundred
decisions, compounded.

Northwind Systems · portfolio valueT–0
01

Your company runs on a decision log.

Take any company with a few years behind it and the line is a record of calls that were made: the product that launched, the migration that ran one more quarter, the pricing change that went live, the new way of working that stuck.

02

Each call opens a fork.

At the moment of the decision both futures are live: F1, where the product ships in March, and F0, where you carry on as you are and ship in September. The difference in outcome between them is the impact of the decision.

03

That impact is several impacts.

The difference shows up on more than one dimension, each measurable on its own: cost, time to market, retention, competitive position, team capacity. Add them up and you have the delta — how much stronger the company stands because this branch was chosen.

04

Now do that forty times a year.

Deltas compound. A run of well-made calls builds a company that outgrows its own plan, and each one looks routine on the day it happens. Year-end fortune is mostly an accumulated hit rate.

05

Your success is one path through the cone.

Raising the hit rate lifts the whole path, decision by decision. Decision quality is the input you can engineer — and the compounding does the rest.

The problem

Humans are unreliable decision makers.

This is the finding cognitive science has reproduced for fifty years, across every profession that has been studied. Facing a call, we take in a great deal that has nothing to do with the impact: who is proposing it, how confident they sound, the months already spent, the customer story from Tuesday. And of the impact itself we usually see one part — the part we happen to know — and treat it as the whole picture.

The impacts that separate F1 from F0

Customers retained past week one€2.4M / year
Platform-team capacity spent9 weeks
Enterprise tier opens later1 quarter
Support load carried by the team€180k / year
Onboarding parity with rivalsdemo-level

What else takes up the room

Seniority of whoever is proposing itweight
Confidence and polish of the pitchweight
Money and months already spentweight
The most recent, most vivid storyweight
Who is in the room and who owes whomweight

Fig. 02 — Illustrative. Attention in a decision meeting is finite. Two impacts get argued, three barely surface, and the rest of the weight goes elsewhere.

One discipline took this finding seriously and built its whole working environment around it.

The precedent

Aviation built for the human it actually had.

The first answer to crashes was better aircraft, and it worked until it ran out. Engines and airframes became reliable, and the accidents kept arriving with pilot error written on them.

Pilot error is an easy verdict and a useless one. The question aviation started asking instead was harder and far more productive: what would stop the next person, in the same seat, on the same night, from doing the same thing? Answering it means changing the situation rather than the person.

So human limits became a design input. Instruments that show the aircraft’s actual state instead of asking a tired person to infer it. Checklists that survive stress. Standard call-outs that make one crew member’s doubt audible to the other. A common language so a first officer can raise a concern to a captain and keep their career. All of it grounded in human factors research and cognitive science, and all of it aimed at a competent crew on a bad night rather than an exceptional pilot on a good one — safety that needs heroics is not safety.

And every incident feeds back, so what one crew learns the hard way, every crew afterwards flies with.

1 in 13.7M
Fatality risk per passenger boarding on commercial flights, 2018–2022 — down roughly an order of magnitude every two decades since the 1970s.
Barnett & Wang, MIT (2018–2022)
1981
First airline-wide Crew Resource Management programme. The lever was how a crew reasons together out loud, in a language every rank can use.
United Airlines, after Portland 1978 & Tenerife 1977
208 sec
From bird strike to touchdown on the Hudson, with 155 people aboard and every one of them walking away.
US Airways 1549, 15 Jan 2009

Those two numbers are the same achievement seen from opposite ends.

One is error made rare across tens of millions of flights. The other is a crew with three and a half minutes and no good options, doing something extraordinary — and afterwards, Sullenberger credited the training, the procedure and the coordination rather than himself.

That is the claim worth borrowing. People are fallible, and the same people, given the right instruments, information and setup, do extraordinary things under the hardest conditions there are.

What that means for a company

You hire good people, and that is worth doing. It is also slow, expensive and only partly in your control — and a company whose results rest on someone exceptional having a good quarter is carrying a risk nobody has priced.

The situation is the part you can build. Good people, the right instruments, and calls that come out well day by day, decision by decision — an engine that keeps running when your best thinker is on holiday, and one that carries a company further than any single brilliant quarter.

The Global Debate Evaluation Standard was built by taking that lesson — aviation’s instruments and loop, plus what cognitive science established about how we weigh information — and applying it to the decisions a society makes together. Reason Based Product Management translates the result into the business world.

The standard is open and documented in full, including the research it rests on: how we think, the framework, and the value register and evidence rules.

The method

Bring clarity into your decisions.

The hardest decisions we make are the ones a society makes together: the stakes are the highest available, and people disagree about values as well as facts. The Global Debate Evaluation Standard was built for that setting — a way to weigh competing arguments in public, on grounds anyone can check.

Reason Based Product Management translates it to the business world.

The standard evaluates anything you can explain. Every argument names exactly one impact — one way in which the two futures differ — and answers the same three questions about it. Arguments become calculable, comparable with each other, and available to learn from a year later.

The default pair is F1, where you act, against F0, where you carry on as you are. The same form takes any pair, so F1 against F2, F3 or any other route on the table is a comparison rather than a fresh debate.

Value — V

How much do we care?

The weight this dimension carries for you: revenue, retention, risk, people, mission. Set once, company-wide, so it holds steady whoever is presenting.

Impact — I

How big is this delta?

The distance between F1 and F0 on this one dimension, in units — €2.4M ARR, nine weeks, four points of churn, 1,200 tickets. Every dimension gets its own line.

Plausibility — P

How well is it grounded?

What the evidence supports today. A modest delta you can back earns more weight than a large one you’re hoping for, and the room can see exactly why.

V × I × P per impact, added up per side: that sum is your signal. And once a decision is written this way, the same form compares it to any other decision on the table.

Decision · Open a self-serve tier for small businessesV × I × P ÷ 10

Reading

Fig. 03 — Two analyses of the same ambition. The first argument on each side is worked through; open any other to see how it scores.

After the decision

You see the moment the arguments stop holding.

Every decision keeps its reasoning, so new information has somewhere to go — and a pivot arrives as a reading rather than a fight.

Project Northwind · tracking against the predictionMonth 9
PREDICTED BAND F0 · CARRY ON AS YOU ARE OUT OF THE BAND ARGUMENTS NO LONGER HOLD COHORT TEST RIVAL SHIPS COST REVISION IMPACT DELIVERED CONFIDENCE IN THE LOAD-BEARING ARGUMENT P 8P 9P 6P 3

Fig. 04 — Each marker is new information going in. Inside the band, the original arguments still carry the project; outside it, they have stopped carrying it.

Stopping a project, or putting another million behind it, usually costs weeks of meetings — because nobody in the room knows where the line sits.

You know where it sits, because the reasoning that produced the decision is still on the table. Put the new information in — the test result, the competitor’s move, the cost that landed differently — and watch the arguments shift. The balance either crosses or it holds, and both answers are useful.

Ahead of the band. The launched product is pulling a bigger delta than you dared assume. Raise the bet now, while the evidence is fresh, with a reason everyone can check.

Slower, mechanism intact. The impact estimate ran optimistic and the evidence still stands. Re-baseline, keep the funding, name the number you’re updating.

Out of the band. The arguments you started with have stopped holding, and the other side is now stronger. Change course this quarter — narrow the scope, or stop and redeploy the team. The case carries the outcome, and the people keep the credit for calling it early.

Months already spent stay in the accounting where they belong, and a course that still holds gets confirmed instead of relitigated.

The compounding part

Your hit rate rises with every decision you record.

Every call carries a prediction, so every outcome measures how well you predicted. Estimates land nearer, confidence starts to mean something, and arguments that used to fill a meeting resolve in minutes.

Predicted vs. actual · gap per decision34 decisions
GAP BETWEEN PREDICTED AND ACTUAL IMPACT DECISION 1 DECISION 34 SPREAD OF YOUR ESTIMATES

Fig. 05 — Illustrative. Calibration improves with feedback: this is the effect Tetlock’s forecasting work found in trained, scored forecasters.

Three things sharpen, and they sharpen for the whole company at once.

Estimates calibrate. Your tenth retention estimate has nine measured outcomes standing behind it.

Confidence gets honest. When your 8s keep landing like 6s, you have a bias with a size attached — and you can correct it on the next call.

Value weights settle. What matters gets argued once, and every later decision inherits the answer.

The learning belongs to the company rather than to the person who happened to be in the room, so one team’s checked prediction sharpens the next team’s estimate. The longer you run it, the higher your hit rate.

What it takes to adopt

Keep everything you have.
Add one framework.

No restructuring, no migration project, no new way of working to learn. It attaches to the moment a call gets made, and everything around that moment stays as it is.

Your setup

Keep your structure

Agile, SAFe, stage-gate or your own hybrid. Teams, org chart and reporting lines stay exactly where they are.

Your people

Keep your roles

Whoever holds the call today holds it tomorrow. Nobody is appointed to a new job to make this work.

Your calendar

Keep your meetings

Existing reviews, steerings and plannings keep their slots and gain a sharper agenda. One page per decision.

What you get for it: more success, because more calls land on the better side. More transparency, because every decision is built from inputs anyone can see and can be re-read when the information changes. And decisions that follow your values, because the weights are yours and they apply the same way to every call in the company.

The return starts in the first session. Decisions with a wide gap between the two futures are the easy ones to identify — a wide gap survives being wrong about any single estimate — so the calls carrying the most consequence come out right straight away. The close ones sharpen after that, and those are the ones that compound: small deltas called correctly, month after month, are what carry you past competitors still deciding on the strength of the presentation.

And who puts it in place?

Argupedia does not consult — our focus is on democratic debate. For adoption inside a company we work with an implementation partner who genuinely knows process, workshops and change.

Who can support you (German) →

A quick read

Could this help you?

Three questions. Answer them the way your last twelve months actually went — if you are already doing this well, it will say so.

a few nearly all
rarely routinely
months days

The longer self-assessment on the business page goes deeper: seven questions, one minute, three numbers at the end — among them the gap between what you believe about your decisions and what your answers show. It is in German.