Section 04 · judge

Confidence is not a feeling.

Every assumption earns its way up an evidence scale. The rating describes the strength of the evidence, not the team's optimism. If we only have what people said, the highest rating is Indicated.


The levels of confidence

How strong is our evidence?

The levels of confidence: four ascending levels of evidence (Assumed, Indicated, Demonstrated, Confirmed at scale) plus Invalidated, a verdict that can land at any level.
Plate 02 The evidence staircase. Four levels, plus one verdict that can land at any level.
L1

Assumed

Belief only. A hypothesis, expert opinion, or untested logic. Every assumption starts here. Honest, not shameful.

L2

Indicated

Signals suggest it may be true. Interviews, surveys, desk analysis, analogous examples. What people say they would do is Indicated at best.

L3

Demonstrated

We observed the behaviour in realistic operating conditions, in one real site. The first level that counts as behavioural evidence. Plausible becomes "we saw it."

L4

Confirmed at scale

Decision-grade evidence. A representative multi-site trial, sustained past novelty, with Data and Finance co-designing the evidence. Suitable for scale or investment decisions.

Invalidated off-ramp at any level

A fair test contradicted the belief. Not failure: a more expensive mistake downstream was avoided. Report it with pride.

The line teams most often blur

Indicated versus Demonstrated. Have we observed actual behaviour in realistic operating conditions? What people told you in a room is never Demonstrated.

The trigger question

Have we observed actual behaviour in realistic operating conditions?

If no

The highest rating available is Indicated: signals only, or what people said. No evidence yet keeps it Assumed.

If yes

Demonstrated for one real site. Confirmed at scale for a representative trial sustained past novelty.

Fig. 12
Observed behaviour in realistic conditions? No Yes no evidencesignals only one real siterepresentative trial Assumed Indicated Demonstrated Confirmed at scale
One question sets the ceiling. Until behaviour is observed in realistic conditions, an assumption tops out at Indicated, however convincing the room was. Observation is the gate to Demonstrated, and a representative trial to Confirmed at scale.

Prioritise the riskiest

Plot each assumption on a 2x2 to find the critical ones.

High risk if wrong, currently only Assumed. These are the critical assumptions. They have the greatest potential to change the direction of the work.

Lower evidence · Assumed
Higher evidence · Demonstrated
High risk if wrong
Test first

Critical assumptions

Could materially change the problem, solution, investment, operating model, or test design.

Test before scaling

Monitor / strengthen for scale

Evidence exists, but broader testing may be needed before commitment.

Low risk if wrong
Only if cheap / near-term

Park or test lightly

Test only if quick, cheap, or needed for an upcoming decision.

Reuse, don't retest

Use as supporting evidence

Do not spend more effort unless conditions change.

When several are critical, ask

What is the next most important assumption that would create significant change if wrong? Test the one most likely to change the problem definition, solution, investment decision, operating model, or test design.

Design the experiment to move confidence

Each test moves one assumption up one level.

From → to

Assumed → Indicated

Learn from people, data, or prior examples.

From → to

Indicated → Demonstrated

Observe behaviour in realistic operating conditions.

From → to

Demonstrated → Confirmed

Decision-grade evidence across representative sites.