← Scrum Master: facilitate, unblock, and improve
09 / 10 · 60 MIN

Experiments, outcomes, and costs

Interpret usage, composition, and effort with explicit assumptions and conclusions bounded by evidence.

Define what you seek to learn

An improvement experiment begins with uncertainty that matters to the service. In the fictional Lume example, the aim is to reduce reconciliation effort without producing incorrect balances. “The team liked the new tool” answers neither part adequately. Before rehearsal, make explicit the expected effect, data to observe, period, observation that would contradict the hypothesis, and who can decide the next step. A stop rule for incorrect balances is a local condition of this exercise, not a universal Scrum rule. If it occurs, retain evidence and follow the agreement; do not change the threshold merely to preserve a favorable conclusion. DORA describes teams able to experiment and adapt solutions with context about business outcomes. That organizational support does not waive authorization for acting in a real environment. An isolated prototype can support bounded learning before a production decision.

Choose the unit and retain groups

A tool has 200 eligible users, forty try it, and twenty complete the task. Observed initial adoption is 40/200=20%; completion among those who tried is 20/40=50%; completion among all eligible users is 20/200=10%. These answer different questions. If 120 identifiers generate 200 opening events, do not write “200 users.” Also investigate non-adopters: interviewing only advocates leaves important barriers outside the conversation. In Aurora, there are initially 81 completions among 90 simple tasks and two among ten complex ones; later there are ten among ten and 36 among 90. Within-group rates improve, but the total falls from 83% to 46% because composition changed. The simple average of 100% and 40% is 70%, but it does not represent the set containing ten and ninety tasks. Keep numerators, denominators, definitions, and sample limits visible.

Distinguish observation, cause, and payback

A change from six returned requests among forty to nine among ninety means 15% to 10%: five percentage points lower and a relative reduction of one third. The absolute count rose. None of these calculations proves the cause. If triage, staffing, and difficulty change together, comparison does not isolate triage’s contribution. A supervised pilot with four completions among five participants describes that pilot; it does not prove independent success for 80% of all users. Include costs too. With 18 initial hours, four saved weekly, and one for maintenance, net savings are three and hypothetical payback takes six weeks. If maintenance exceeds savings, there is no payback in that constant model. Other benefits may exist but need identification. Do not end observation after three easy days if the proposed conclusion includes a funds closing that has not yet been observed.

Reproducible interpretation lab

Copy the complete code below into a file named experiencia.py and run python3 experiencia.py using Python 3.13. It needs no external packages, network, personal data, or platform access. The program uses exact fractions for fifteen checks: before-and-after rates, percentage points, relative reduction, adoption, conditional and overall completion, composition, averages, and payback. It rejects a zero denominator instead of presenting it as 0%. Before running it, calculate results manually and write a sentence about what each calculation does not establish. Then compare with expected and actual results. In a new exercise file, change values and expectations consistently; a failed assertion indicates disagreement with the expected result, not failure of a team. Produce an interpretation of Aurora, Lume’s cost assumptions, and an observation plan. The fifteen checks were executed with CPython 3.13.1. No experiment with real teams, causal proof, statistical significance test, or observed field savings occurred.

"""Original fictional-data arithmetic; no field experiment or causal inference."""
from fractions import Fraction
from pathlib import Path
import hashlib
import json
import platform

checks = []
def check(name, actual, expected):
 assert actual == expected, (name, actual, expected)
 checks.append(dict(name=name, actual=actual, expected=expected, passed=True))

def rate(successes, total):
 if total <= 0 or not 0 <= successes <= total:
 raise ValueError('Counts require 0 <= successes <= total and total > 0')
 return Fraction(successes, total)

def payback(initial_hours, saved_weekly, maintenance_weekly):
 assert initial_hours > 0 and saved_weekly >= 0 and maintenance_weekly >= 0
 net = saved_weekly - maintenance_weekly
 return None if net <= 0 else Fraction(initial_hours, net)

before, after = rate(6, 40), rate(9, 90)
check('before_return_rate', before, Fraction(3, 20))
check('after_return_rate', after, Fraction(1, 10))
check('percentage_point_drop', 100 * (before-after), Fraction(5))
check('relative_drop', (before-after)/before, Fraction(1, 3))
check('initial_adoption', rate(40, 200), Fraction(1, 5))
check('completion_among_triers', rate(20, 40), Fraction(1, 2))
check('completion_among_eligible', rate(20, 200), Fraction(1, 10))
check('pooled_before', rate(81+2, 90+10), Fraction(83, 100))
check('pooled_after', rate(10+36, 10+90), Fraction(46, 100))
check('both_groups_improved', rate(10, 10)>rate(81, 90) and rate(36, 90)>rate(2, 10), True)
check('unweighted_mean_is_not_pooled', (rate(10, 10)+rate(36, 90))/2, Fraction(7, 10))
check('six_week_payback', payback(18, 4, 1), Fraction(6))
check('no_payback_negative_net', payback(12, 2, 3), None)
check('no_payback_zero_net', payback(12, 3, 3), None)
try:
 rate(0, 0)
except ValueError:
 zero_rejected = True
else:
 zero_rejected = False
check('zero_denominator_is_not_zero_percent', zero_rejected, True)
print(json.dumps(dict(groups=len(checks), checks=checks, python=platform.python_version,
 scope='Synthetic counts and constant effort assumptions. No statistical significance, causal proof, actual savings, personnel assessment or production change.',
 scriptSha256=hashlib.sha256(Path(__file__).read_bytes).hexdigest), default=str, ensure_ascii=False, indent=2))
IN PRACTICE

6/40=15% and 9/90=10%: the rate fell five percentage points despite the larger count. The cause requires other evidence.

Common pitfalls

Events treated as people; unweighted mean treated as total; assisted pilot treated as independence; omitted maintenance; correct calculation treated as causal proof.

Related topics: Improvement hypotheses · User feedback · Metrics and costs

Take this idea with you

A reproducible calculation supports evidence discussion; conclusions must respect the population, conditions, and observed assumptions.

Create account

Reference: Team experimentation · Scrum Guide November2020; Kanban Guide May2025; EBM May2024; primary guidance reviewed 2026-09-30