Measurement and falsifiability

RDD should be measured by where disagreement appears.

RDD is not a belief system. After a team tries it, measure whether consequential ambiguity moved earlier and whether less misunderstood behavior crossed the delivery baseline.

Do not invent benchmarks. Start with a ledger that makes timing, evidence, and post-baseline churn visible.

What to measure after trying RDD

Track where ambiguity moves.

MetricWhat it revealsWhen to measure
Time from initial intent to executable stakeholder feedbackWhether reality arrives early enough to affect commitment.During Build to Learn.
Significant disagreements surfaced during Build to LearnWhether candidate reality is exposing consequential ambiguity.Each Reality Loop.
Functionality discarded before production engineeringWhether the team is avoiding unnecessary committed scope.Before Freeze Gate.
Post-baseline behavioral scope churnWhether misunderstood behavior is still crossing into delivery.During Build to Deliver.
UAT defects caused by misunderstood intentWhether late validation is still carrying discovery work.UAT and release review.
Late change requests caused by ambiguity versus new needsWhether change is avoidable misunderstanding or genuine new information.Change-control review.
Reality Checks passed, failed, or deferredWhether accepted behavior survives technical and operational reality.Before and after baseline.
Production scope backed by accepted executable evidenceWhether delivery is tied to experienced behavior rather than prose alone.Baseline and release planning.

The practical standard

RDD should be measured by where disagreement appears and how much ambiguity crosses the delivery baseline.

If a team runs Reality Loops and still discovers the same behavioral misunderstandings in UAT, the method did not do enough. If the loops surface disputes early, discard weak scope, and give engineering a clearer baseline, the method has done useful work even before any budget effect is claimed.

What not to measure alone

Do not judge RDD only by how quickly a candidate was generated. Speed is useful only when it creates better evidence. A fast prototype that carries false confidence into production is not an RDD success.

The stronger test is whether the team can explain why each committed behavior exists, which loop accepted it, which alternatives were rejected, and which Reality Checks still limit confidence.

Start with one measurable loop

Use a first Reality Loop to establish the initial uncertainty, evidence needed, and decision record.