Plan evidence before observing the result
A pilot should answer a specific uncertainty: does the solution reduce manual work without increasing failures, or support close volume? Define eligible groups, period, metrics, acceptance criteria and stop-or-investigate conditions beforehand. If the team chooses exclusions after seeing unfavorable results, reporting can favor the solution without reflecting the need. Also record who decides and which information supports expansion, correction or deferral. Exercise criteria are fictional and do not replace organizational standards. A favorable small-pilot result may justify the next collection without justifying expansion to every fund, volume and operating period.
See what changed in the population
The synthetic baseline has eighty simple requests with 72 successes and twenty complex requests with twelve successes. The aggregate is 84%. The pilot has twenty simple requests with nineteen successes and eighty complex requests with 52 successes. The aggregate falls to 71%, although simple requests improve from 90% to 95% and complex requests from 60% to 65%. Composition changed substantially: the pilot contains more difficult cases. Declaring solution deterioration solely from the aggregate would ignore that difference. It would also be wrong to conclude every result is satisfactory merely because each group improved five percentage points. The business criterion may remain unmet, especially in the higher-risk group.
Standardize composition without inventing causality
Applying baseline weights, 80% simple and 20% complex, to pilot rates gives 89%. The standardized difference from 84% is five percentage points. This calculation answers a descriptive question: what rate follows from observed rates under a common composition? It does not prove that the solution caused improvement. The team, period, classification rule or difficulty within each group may have changed. Document those hypotheses and avoid calling a simple before-and-after comparison an experimental control. If stronger causal conclusions are needed, plan an appropriate evaluation with specialists and operational feasibility instead of increasing decimal precision on the same limited dataset.
Combine criteria without hiding a gap
The readiness exercise has four mandatory conditions: functional behavior, recovery, evidence coverage and an operational owner. Three are confirmed but coverage is missing. The explicit rule requires all four, so readiness remains false. A 75% average does not authorize replacing the missing condition. In actual work, distinguish demonstrated failure, unknown information and a justified not-applicable criterion. Each state calls for a different action. The team may prepare a conditional recommendation, but acceptance of the condition belongs to the defined authority. Link criteria to tested version and scope so an old result is not reused as proof of an unevaluated configuration.
Communicate the decision supported by data
Prepare a summary with group results, aggregate, coverage, differences from baseline and relevant uncertainties. For the sponsor, explain the recommendation’s consequence and next decision point. For APS, add effects on exceptions, manual intervention, capacity and resumption. If proposing phased expansion, identify which hypothesis each phase will test and which observation can halt expansion. Retain results contradicting expectations as well. Automated calculation validation shows that the program applied fixture rules; it does not establish that actual data are complete, the selected criterion represents value or the organization authorized advancement. Those elements need their own evidence.
# Synthetic standardized pilot result using baseline group weights.
standardized = 0.8 * 0.95 + 0.2 * 0.65 # 0.89
# This is descriptive adjustment, not causal proof.
# Readiness additionally requires every agreed mandatory condition.The aggregate falls from 84% to 71%, but both groups improve five percentage points; under common composition the pilot produces 89%.
Common pitfalls
Composition change as causal effect; averaging mandatory criteria; redefining exclusions after results; extrapolating five observations to production.
Related topics: Needs and business case · Metrics and acceptance criteria · Operations and benefit realization
A useful comparison retains population, definition, context and inference limits.
Reference: Ten Guidelines for Successful Benefits Realization · Five-domain ECO / verified 2026-10-01