A consistent table can express the wrong decision
The lab creates a variant where R2 changes from REVIEW to AUTO. The table still has no gaps or overlaps: every valid input receives one output. However, the separately established business example requires REVIEW for a confirmed transaction of 150000 cents, even with 120 minutes remaining. The variant fails that example. This distinguishes checking model quality from validating suitability for the need. If a test calculates its expected answer from the changed rule itself, both can agree on the same error. Elicit expected examples from authorized people, include conditions that challenge the proposal and retain the origin of the expected result. One user’s opinion can be elicitation evidence without being sufficient approval.
Compare versions without turning differences into benefits
Proposal v2 extends the review window from 15 to 20 minutes before close. In a fictional set of 24 examples, three change from AUTO to REVIEW: confirmed transactions of 50000 cents with 16, 18 and 20 minutes remaining. The remaining examples keep their output. The comparison identifies decisions to discuss; it does not prove v2 is better. The organization may want less automatic processing near close, but the new review can create a queue without sufficient capacity. Ask which problem the change should reduce, who receives the additional work and which evidence will evaluate the outcome. The three examples do not estimate affected daily volume: they were selected to reveal boundary behavior rather than sampled from real traffic.
Distinguish effectivity, deployment and event date
The fixture uses organizational dates and intervals with inclusive starts and exclusive ends. V1 applies from October 1 until October 15; v2 starts on the 15th. Querying the 14th returns v1 and querying the 15th returns v2. Two versions valid on the same day create ambiguity that the script rejects. Production would require much more clarification: time zone, the event determining the version, handling of requests received before the change and reprocessed afterward, and application compatibility. Installing new code does not automatically decide which policy applies to an earlier event. The analyst should make these choices explicit and relate rule, approval, effective date, implementation and tests. The lab does not model distributed clocks or actual replay.
Prepare operations and observe the relevant outcome
A dashboard says 99% of decisions produced an output but excludes inputs rejected for missing reconciliation. Before accepting that measure, define its denominator, exclusions and exception handling. For the v2 change, track volume routed to review, queue age, close completion and confirmed routing errors. The definition should distinguish an intended policy change from an engine failure. Retain the version and data needed to explain the decision under agreed access and retention controls; do not collect sensitive fields unnecessarily. APS needs to know who acts when no rule applies, when data is incomplete and when the queue cannot be handled in time. A technically resolved state does not establish completion of the business process.
Organize review and retain evidence limits
Prepare a decision package containing the need, changed conditions, changed and unchanged examples, confirmed expected results, operational impact and approval authority. Structural checks and the lab’s two runs support technical review; they do not replace stakeholder validation or production authorization. If a new reconciliation state appears, review the domain, rules, interfaces, measures and examples instead of silently mapping it to confirmed. Retain the earlier baseline to explain historical decisions and distinguish defect correction from an intentional policy change. The conclusion should state what was demonstrated, in which version, for which input domain and which human decisions remain.
python3 content/labs/cbap-decision-rules/run.py --output /tmp/cbap-rules.json
# Inspect gapWitness, conflictWitness and changedExamples.V2 changes 3 of 24 examples. An incorrect R2 variant passes structural checks but fails the expected REVIEW outcome for 150000 cents.
Common pitfalls
Deriving expected output from the rule under test; treating synthetic differences as actual frequency; confusing deployment date with effectivity; hiding rejections from the denominator.
Related topics: Requirements architecture · Validation and acceptance · Rule life cycle
Verification establishes bounded model properties; validation connects the rule to the need, examples and consequences in work.
Reference: CBAP competencies · CBAP six-knowledge-area blueprint, May 2026 handbook