1. Separate stopping, holding and expanding
A fictional release initially receives a limited set of instructions. The agreed rule stops expansion when there are at least 100 completed instructions and five or more failures; permits proposing expansion only after 100 completions, 30 minutes of observation and fewer than five failures; every other state retains current exposure. Final authorization belongs to the service owner. At minute 18, there are 120 completions and two failures. No stopping condition is established, but minimum duration has not elapsed: retain exposure. At minute 35, four failures among 140 completions satisfy these numerical and temporal conditions, permitting an expansion proposal. They do not mean automatic authorization. This is only a teaching contract, not a threshold derived from statistical probability. On an actual project, owners must justify thresholds and duration against risk and work type.
2. Retain the version associated with each result
Within a defined period, the old version completed 900 instructions, nine failed; the new version completed 100, eight failed. Overall there are 17 failures in 1,000, or 1.7%. However, the new version has 8% failures while the old version has 1%. If the new-version criterion allows at most 2%, the aggregate does not establish compliance: the new version fails the observed criterion. The contrast alone does not prove which component caused failure. Populations may differ in relevant characteristics. The recommendation should retain both conclusions: the stated criterion failed, and causal explanation needs investigation. Keep effective version, period and unit of work with each result; a label added only on the final dashboard may not recover this information.
3. Check absolute limits when groups share dependencies
The new group and comparison group use the same database. During a window both have 8% failures. The rule requires both a rate below 5% in each group and a between-group difference of at most one percentage point. The observed difference is zero and satisfies the second condition; both rates violate the first. This is not indeterminate merely because the contrast is small: the absolute condition is demonstrably out of bounds. Records also show database saturation, but do not yet identify the load source. The planned action is to suspend expansion and investigate under agreed operational procedures. Do not automatically attribute all saturation to the new version. A shared dependency can make the comparison group experience the same effect; the problem remains even when their results converge.
4. Observe the work unit through its relevant outcome
A reconciliation batch has three steps: reading, calculation and publication. Evaluation aims to compare a complete batch executed by one version. Current routing chooses a different group at each step. Even with version labels on every record, a batch passing through both groups is not an end-to-end outcome of one version. The team needs to retain batch assignment or formulate a different evaluation question explicitly covering mixed execution. In another trial, 60 batches started; at minute eight, 42 have finished and 18 remain active. The requirement permits 12 minutes per batch. Those 42 completions do not establish that all 60 met the deadline, and the 18 active batches are not yet demonstrated failures. Retain each batch’s start time and observe its outcome at the relevant limit.
5. Measure exposure in the unit relevant to the business
A selection contains 50 of 1,000 clients, or 5%. Those clients nevertheless account for 400 of the window’s 1,000 instructions, or 40%. There is no contradiction: clients and instructions are different denominators. If authorization limits exposure to 10% of instructions, the selection exceeds that limit despite involving few clients. Nor can 40% of instructions automatically be converted to 40% of financial value. For a decision, obtain the attribute used by the limit and retain the window definition. Selecting only low-volume clients may reduce initial exposure while omitting the most demanding batches. Document both effects without presenting a low-risk initial choice as representative of every client. Expansion should add coverage deliberately.
6. Match the reporting window to the question
The new version starts at 10:00. Between 10:00 and 10:05 it completes ten instructions, all failed. A dashboard covering 09:00 to 10:05 mixes them with 990 earlier failure-free completions and displays 1%. That overall calculation is correct for the displayed set, but does not answer the new version’s failure rate since activation: its observed rate is 100%. Changing the report title is not enough. Query and data must select the defined population and period. Before the meeting, write the complete question: version, start, end, unit and eligible outcomes. For accumulated metrics, also confirm that you compare increments from the same period. A screenshot without these elements may be insufficient to reproduce the conclusion.
7. Distinguish association, cause and the permitted decision
Before introducing new code, the team runs the same version in two groups. One receives only small files and the other large files. It observes 1% and 7% failure rates. The difference cannot be attributed to a code change that has not happened. The trial exposes a population, configuration, measurement or other difference to investigate. It does not automatically identify which explains the result. When the new version arrives, retaining the same assignment may continue to confound comparison. The analyst should explain that the planned comparison does not isolate the version effect and propose a suitable design with technical owners. Equal trial results would not establish that all future groups will be equivalent either. The requirement is to keep conclusions proportionate to evidence, not turn every contrast into causal certainty.
8. Exercise: prepare the expansion recommendation
Write a recommendation before reading the solution. Authorization requires at least 100 completed batches, 30 observed minutes, fewer than five failures and exposure to at most 10% of the window’s instructions. At minute 35, the new version has 120 completions and three failures; selected clients account for 18% of instructions. The team presents only “3/120, ready to expand”. Solution: count, time and failure limit meet the supplied conditions. Exposure violates the 10% limit, so the combined conditions do not permit proposing expansion. Communicate observed exposure and apply the action specified in the agreement without pretending that correcting the denominator removed exposure that already occurred. An authorized limit revision is a future decision, not proof of past compliance. In the English meeting, you might say “the outcome conditions are met, but exposure exceeds the agreed limit” and identify who decides the next step.
Exercise files
Original Python exercise with editable data, instructions and worked reasoning on intersections, availability, histograms and expansion decisions. Requires Python 3.10 or later.
A release has 8% failures, but its aggregate with the old version shows 1.7%. The release decision uses results from the defined scope.
Common pitfalls
Confusing absence of a stop condition with authorization; mixing versions or periods; relying only on group differences; counting steps as complete batches; translating client percentage into exposed value.
Related topics: Acceptance criteria and observation · Impact analysis and deployment decisions
Expansion must meet all agreed conditions and respect defined authority; one favorable indicator does not replace the complete decision.
References
- PMI-PBA Examination Content Outline · Public five-domain ECO; body copyright 2013/back cover 2017; inspected 2026-10-09
- Canarying Releases · Public chapter inspected 2026-10-09; contextual engineering reference
- Monitoring Distributed Systems · Public book chapter inspected 2026-10-09; contextual engineering reference