← Problem Management: investigate and prevent
08 / 8 · 60 MIN

Hypotheses, regression, and closure

Build an evidence chain between hypothesis, change, and effectiveness using a bounded resource-leak model.

Separate recovery and causal attribution

The fixture changes three variables together: version v1 to v2, load from 1000 to 200 requests per minute, and cache from warm to empty. Observed failures go from 20 to zero. Reporting improvement is correct, but assigning it exclusively to version is not. Reduced load may avoid reaching the limit; resetting cache may postpone the symptom. Record every change and write hypotheses producing different predictions. In an authorized environment, compare sufficiently similar conditions and observe the expected mechanism. A canary also needs characterized populations: routing only simple operations to the new version does not by itself create a useful comparison with complex operations in the control.

Read accumulation before failure appears

The Python model represents a resource counter with capacity ten. The broken variant retains one resource every ten successful operations. After fifty operations, there are fifty successes and five retained resources: a green test hides movement toward the limit. Across 500 attempts, the variant completes 100 and rejects 400, ending with ten retained resources. The repaired variant releases each resource, completes 500, ends at zero, and has a peak of one. These results follow the explicit rules of the sequential model. They do not measure a JVM, JDBC pool, or WebSphere, exercise concurrency, or define a recommended size for real pools.

Design a regression that distinguishes versions

A useful regression fails through the relevant mechanism in the earlier version and passes for the expected reason in the proposed change, where that comparison is safe and feasible. In the model, fifty operations without retention measurement pass both versions and do not distinguish the defect. A longer run and the counter reveal the difference. For a real product, identify acquisition and release paths, cancellations, exceptions, and workload conditions requiring exercise. Define stop and recovery criteria before execution. Do not transfer 500 as a universal standard. Retain version, configuration, inputs, and results, including counterexamples; if the defect cannot be reproduced, state the limit and investigate missing conditions.

Measure workaround cost and scope

Resetting the counter after every fifty successes avoids reaching the model limit, but retention returns between resets. This measure changes exposure to exhaustion rather than the release rule. In a fictional APS service, a periodic restart may reduce incidents while increasing manual interventions and error risk during the closing window. Report availability, recurring effort, and intervention effects separately. Define use conditions, failure signals, and an owner for workaround review. Eliminating the cause needs its own evidence. If correction is delayed, a decision to continue with mitigation should expose cost and residual risk to the people authorized to accept them. Keep that decision separate from the technical claim of removal.

Link each action to a verifiable result

An approved commit establishes a reviewed change; an installed change establishes another stage. Neither reference alone demonstrates that the mechanism stopped occurring. For each action, record owner, artifact, verification conditions, expected result, and obtained evidence. If the team has not reproduced the leak or measured retention, write effectiveness pending and identify the next exercise. In the postmortem, separate observed facts, hypotheses, and decisions. This structure allows correcting an explanation without erasing events. When another application uses the same library, examine version and cancellation path before generalizing the fix. Sharing a document is a learning action; it does not confirm effectiveness in every consumer.

Define the observation needed for a decision

Zero failures in zero runs has no defined rate; zero in one hundred runs gives an observed zero rate limited to exercised conditions. Acceptance should include relevant exposure, business outcome, and integrity alongside HTTP codes. A canary can return 200 while duplicating entries, so transport and effect need distinct criteria. Before handover to RUN, agree which signals will be followed, who reconciles unknowns, and when effectiveness will be reviewed. If the window ends before sufficient exposure, present observations and the remaining gap. Operational recovery, administrative closure, and cause eliminated are assertions requiring different evidence; do not merge them to meet a committee date.

python3 content/labs/problem-evidence/run.py
# broken, 50 attempts: 50 success, 0 failed, 5 retained
# broken, 500 attempts: 100 success, 400 failed, 10 retained
# repaired, 500 attempts: 500 success, 0 failed, 0 retained
# Scope: sequential counter model; no vendor runtime executed.
IN PRACTICE

A vendor delivers a client fix before a prolonged batch. The team prepares an authorized comparison of error paths, observes retained resources, and separates patch delivery from demonstrated effectiveness.

Common pitfalls

Accepting a test that also passes broken code, calling a reset a fix, confusing a commit with effectiveness, or presenting a sequential model as real middleware validation.

Related topics: Change Management · Middleware · Application Production Support

Take this idea with you

The verification chain links conditions, mechanism, change, and outcome. A recovered service may still require investigation and an explicit residual-risk decision.

Create account

Reference: Conduct post-incident analysis · Problem management practices 2026-09; ServiceNow Brazil examples with scoped plugins and properties