← Product Owner: discover, decide, and create value
10 / 10 · 60 MIN

Metrics, identities, and experiments

Calculate outcomes with explicit units, check data, and separate observation from causal attribution.

Name the population and unit

In the fictional Iris pilot, sixty people are eligible in the same window. Twenty-four start at least one task and eighteen complete at least one. Adoption defined as starting is 24/60=40%; completion among starters is 18/24=75%; completion coverage among eligible people is 18/60=30%. None of these measures should be presented without its denominator. One person starts and completes two operations, yielding 25 started and 19 completed operations. Operation completion is 19/25=76%, different from person completion. The export contains 51 rows, seven exactly repeating earlier events. There are 44 unique events, not 44 users. Establish the meaning of event_id, user_id, and operation_id before using dashboard counts. This laboratory’s values are invented data, not bank statistics or evidence that a feature causes adoption or creates value.

Check collection before concluding

The exercise contract allows repetitions only when identity and content match. Two rows sharing an event_id with different states should produce an error rather than silently choosing the last row. The script also rejects a person outside the declared population, unknown type, blank identifier, and an operation assigned to two people. The synthetic window is complete and contains all starts, so a completion without a start is rejected. A real export with a partial window or delayed events would require a different collection and reconciliation contract. The code does not attempt to solve those cases. It does not convert blank duration into zero: for 10, blank, and 20, the mean of the two observations is fifteen and coverage is two thirds. That partial mean does not automatically describe unmeasured tasks. Retain the gap and investigate whether event loss depends on outcome or user type.

Interpret benefit and experiment

In the Delta case, the same task set requires 180 operator minutes and 42 APS minutes before; afterward it requires 140 and 60. Total reduction is 222−200=22 minutes, not forty. Another record shows four errors in 80 operations and nine in 90: 5% and 10%, a five-percentage-point increase and a doubled rate. These calculations do not isolate causes. For an experiment, define the hypothesis, primary metric, population, window, collection, and stop conditions in advance with owners able to act. GOV.UK technical A/B guidance helps prepare design and QA, but its internal register and tools are not banking requirements. A team using A in the morning and B in the afternoon with different tasks has not created randomization by using two variants. Label subsequent analyses as exploratory. Do not compare thirty-day retention for a cohort observed for only seven days as though its window were complete.

Reproducible Iris laboratory

Copy the complete code shown in this lesson into a local file and run it with Python 3.13. It uses only the standard library and in-memory generated data; no platform or banking-system access is needed. Before running, predict people, operations, and rates. Output includes twenty checks, counts, and a missing-duration example. Then change a repeat to have conflicting content and observe the rejection exercised by the test code. Also try an empty population: the rate is undefined, represented by null in JSON, rather than zero. Explain in writing why 76% per operation does not answer the question about people and why the partial mean does not prove complete coverage. The code was executed locally to check these synthetic data. It performs no statistical test, establishes no causality, validates no real pipeline, and does not replace experiment design with appropriate analytical support.

"""Original synthetic teaching model. No files, network, or real user data.
PT: copiar este código completo e executar com Python 3.13.
EN: copy this complete script and run with Python 3.13.
Rates describe this fixture; they establish no causal or statistical inference.
"""
import csv
import io
import json
import math

FIELDS = ['event_id', 'user_id', 'operation_id', 'kind']
ELIGIBLE = {f'u{i:02}' for i in range(1, 61)}

def export_fixture:
 rows = []
 for i in range(1, 25):
 rows.append([f's{i:02}', f'u{i:02}', f'op{i:02}', 'start'])
 for i in range(1, 19):
 rows.append([f'c{i:02}', f'u{i:02}', f'op{i:02}', 'complete'])
 rows += [['s25', 'u01', 'op25', 'start'],
 ['c25', 'u01', 'op25', 'complete']]
 rows += [rows[i][:] for i in [0, 1, 2, 3, 24, 25, 26]]
 stream = io.StringIO
 writer = csv.writer(stream)
 writer.writerow(FIELDS)
 writer.writerows(rows)
 return stream.getvalue

def parse(text):
 reader = csv.DictReader(io.StringIO(text))
 if reader.fieldnames!= FIELDS:
 raise ValueError('unexpected columns')
 return list(reader)

def ratio(numerator, denominator):
 return None if denominator == 0 else numerator / denominator

def summarize(rows, eligible):
 events = {}
 for row in rows:
 if set(row)!= set(FIELDS) or any(not row[k] for k in FIELDS):
 raise ValueError('missing or extra field')
 if row['user_id'] not in eligible:
 raise ValueError('user outside declared population')
 if row['kind'] not in {'start', 'complete'}:
 raise ValueError('unknown event kind')
 key = row['event_id']
 if key in events and events[key]!= row:
 raise ValueError('conflicting event identity')
 events[key] = row.copy
 starts, completes, owners = set, set, {}
 for row in events.values:
 op, user = row['operation_id'], row['user_id']
 if op in owners and owners[op]!= user:
 raise ValueError('conflicting operation owner')
 owners[op] = user
 (starts if row['kind'] == 'start' else completes).add(op)
 if not completes <= starts:
 raise ValueError('completion without start in declared complete window')
 starters = {owners[op] for op in starts}
 completers = {owners[op] for op in completes}
 return dict(rows=len(rows), unique_events=len(events),
 duplicate_rows=len(rows)-len(events),
 starters=len(starters), completers=len(completers),
 started_operations=len(starts), completed_operations=len(completes),
 adoption=ratio(len(starters), len(eligible)),
 completion_among_starters=ratio(len(completers), len(starters)),
 completion_among_eligible=ratio(len(completers), len(eligible)),
 operation_completion=ratio(len(completes), len(starts)))

def durations(values):
 observed = [float(v) for v in values if v!= '']
 if any(v < 0 or not math.isfinite(v) for v in observed):
 raise ValueError('invalid observed duration')
 return dict(observed=len(observed), missing=len(values)-len(observed),
 mean=ratio(sum(observed), len(observed)),
 coverage=ratio(len(observed), len(values)))

checks = []
def check(name, condition):
 if not condition:
 raise AssertionError(name)
 checks.append(name)

def rejects(name, action):
 try:
 action
 except ValueError:
 checks.append(name)
 else:
 raise AssertionError(name)

rows = parse(export_fixture)
r = summarize(rows, ELIGIBLE)
check('row identities', (r['rows'], r['unique_events'], r['duplicate_rows']) == (51, 44, 7))
check('people', (r['starters'], r['completers']) == (24, 18))
check('operations', (r['started_operations'], r['completed_operations']) == (25, 19))
check('adoption', r['adoption'] ==.4)
check('completion among starters', r['completion_among_starters'] ==.75)
check('completion among eligible', r['completion_among_eligible'] ==.3)
check('operation completion', r['operation_completion'] ==.76)
check('input order is irrelevant', summarize(list(reversed(rows)), ELIGIBLE) == r)
extra = summarize(rows + [rows[0].copy], ELIGIBLE)
check('extra duplicate changes only raw and duplicate counts',
 extra == {**r, 'rows': 52, 'duplicate_rows': 8})
rejects('conflicting event', lambda: summarize(rows + [{**rows[0], 'kind':'complete'}], ELIGIBLE))
rejects('unknown user', lambda: summarize([{**rows[0], 'user_id':'u99'}], ELIGIBLE))
rejects('unknown kind', lambda: summarize([{**rows[0], 'kind':'clicked'}], ELIGIBLE))
rejects('orphan completion', lambda: summarize([rows[24]], ELIGIBLE))
rejects('operation owner mismatch', lambda: summarize(rows + [dict(event_id='x',user_id='u02',operation_id='op01',kind='complete')], ELIGIBLE))
rejects('blank identity', lambda: summarize([{**rows[0], 'event_id':''}], ELIGIBLE))
check('empty population is undefined not zero', summarize([], set)['adoption'] is None)
d = durations(['10', '', '20'])
check('missing differs from zero', d == dict(observed=2,missing=1,mean=15,coverage=2/3))
rejects('negative duration', lambda: durations(['-1']))
rejects('nonfinite duration', lambda: durations(['nan']))
check('net effort includes support', (180+42)-(140+60) == 22)
print(json.dumps(dict(python_scope='3.13 standard library',checks_passed=len(checks),
 checks=checks,fixture=r,duration_example=d,
 limitation='Synthetic arithmetic only; no real users, causal inference or production experiment.'), indent=2))
IN PRACTICE

24/60 measures adoption; 18/24 measures completion among starters; 19/25 measures operation completion. These are different questions about the same sample.

Common pitfalls

Rows as people; duplication as adoption; missing as zero; before-and-after as cause; a metric chosen after results as the original hypothesis.

Related topics: Product metrics · Data quality · Experiments and decisions

Take this idea with you

Confidence in an outcome begins with measure definition and data quality and ends with a conclusion proportionate to evidence.

Create account

Reference: How to set performance metrics for your service · Scrum Guide November2020; EBM May2024; primary product practice reviewed 2026-09-30