← Operational risk: fundamentals, controls, and decisions
07 / 10 · 60 MIN

Control evidence: population and effectiveness

Build a traceable conclusion from criteria, observations, and coverage gaps.

Start with the claim needing evidence

In a fictional funds project, the team wants to claim that unapproved changes are blocked before execution. Write that claim before collecting documents. Define the action, system, period, population, and expected outcome. Training attendance shows participation; a procedure shows design; a denied request under known conditions observes behavior. These elements can complement one another. Section 2.4 of NIST SP 800-53A Revision 5 distinguishes examination, interview, and testing. This lesson applies that distinction to an original exercise without turning its code into a complete NIST assessment program or a banking audit.

Define the population before selection

The inventory contains two hundred changes: one hundred eighty normal and twenty emergency. Selection contains twenty normal changes, all satisfactory. You observed ten percent of the population and no emergency case. The result can be useful but does not demonstrate the mechanism the shift uses during an emergency. Choosing only the easiest records introduces a limitation that belongs in reporting. Justify selection, mechanism differences, criticality, and period. The example calculates no statistical confidence and prescribes no universal twenty-record sample size. An incomplete population also compromises selection, even when every element of that population was checked.

Reconcile identity and provenance

Four rows do not prove four distinct executions. If you expect a,b,c,d and receive a,b,b,c, d is missing and b appears twice. Retain references to reconstruct origin and distinguish identical repetition from conflict. If the same execution has pass and fail, row order does not decide which outcome is valid. Look for a version, rerun, or transformation explaining the divergence. Until then, keep the exception visible. The local model identifies missing items, unexpected items, copies, and conflicts; it does not automatically select a favorable outcome or delete records from a real system.

Separate observed, failed, and unknown

One hundred controls were scheduled: eighty-five have pass, five fail, and ten lack evidence. Known-result coverage is ninety percent; demonstrated pass fraction across scheduled controls is eighty-five. Dividing eighty-five by ninety answers a different question: outcome among observed records. It may be shown if the population and ten unknowns remain visible. Do not silently fill absent results with yesterday’s state. Do not declare one hundred failures to compensate for uncertainty either. Show what is known, what is missing, and who will recover the observation. Unknown status is information about evidence rather than a fabricated execution result.

Reassess after a change

An approval rule changed after the test. The control name stayed the same, but that does not demonstrate unchanged behavior. Identify the affected part and the additional evidence needed. An external assessment can contribute if its scope, period, and configuration are relevant; a new presentation date does not update the test. In the case of two accounts owned by one person, installing a fix is a technical milestone. If both can still approve, intended independence remains unproven. Keep delivery and effectiveness as distinct states with an owner and the next acceptance criterion.

Execute and explain the limits

Save the following code as run.py and execute python3 run.py --output evidence.json. It uses only the standard library and invented data. Examine each check before interpreting passed: that field counts deterministic comparisons rather than approved banking controls. The program works on in-memory lists and writes a local file; it neither reads customer data nor contacts services. Then explain which evidence would be missing for assessment in an authorized environment. Deliver a matrix containing objective, population, selection, results, gaps, and action. Human review and observation of the real mechanism remain separate activities.

"""Original fictional evidence model; no bank data, capital or compliance assessment.
Run: python3 run.py --output evidence.json
"""
import argparse
import hashlib
import json
import platform
from fractions import Fraction
from pathlib import Path


def ratio(numerator, denominator):
 if denominator == 0:
 return None
 return float(Fraction(numerator, denominator) * 100)


def reconcile(expected, observations):
 by_id = {}
 conflicts = []
 duplicates = 0
 for row in observations:
 key = row['id']
 if key in by_id:
 if by_id[key]!= row:
 conflicts.append(key)
 else:
 duplicates += 1
 else:
 by_id[key] = row
 return dict(missing=sorted(set(expected)-set(by_id)),
 unexpected=sorted(set(by_id)-set(expected)),
 conflicts=sorted(set(conflicts)), duplicate_copies=duplicates)


def loss_summary(events, recoveries):
 # One recognized event per stable ID; identical copies collapse, conflicts fail.
 indexed = {}
 for row in events:
 if row['id'] in indexed and indexed[row['id']]!= row:
 raise ValueError('conflicting event version')
 indexed[row['id']] = row
 recovery_ids = set
 received = 0
 for row in recoveries:
 if row['id'] in recovery_ids:
 raise ValueError('duplicate recovery identifier')
 recovery_ids.add(row['id'])
 if row['event'] not in indexed:
 raise ValueError('unmatched recovery')
 if row['status'] == 'received':
 received += row['amount']
 gross = sum(r['amount'] for r in indexed.values if r['kind'] == 'loss')
 return dict(events=len(indexed), gross=gross, received=received, net=gross-received)


def run:
 checks = []
 def check(name, actual, expected):
 assert actual == expected, (name, actual, expected)
 checks.append(dict(name=name, actual=actual, expected=expected, passed=True))
 check('selected normal records pass', ratio(20,20), 100)
 check('normal stratum sample coverage', ratio(20,180), float(Fraction(100,9)))
 check('emergency stratum sample coverage', ratio(0,20), 0)
 check('population sample coverage', ratio(20,200), 10)
 check('no observed executions is not perfect effectiveness', ratio(0,0), None)
 statuses = ['pass']*85 + ['fail']*5 + ['unknown']*10
 check('complete control evidence coverage', ratio(sum(s!='unknown' for s in statuses),len(statuses)),90)
 check('demonstrated pass fraction of scheduled population', ratio(statuses.count('pass'),len(statuses)),85)
 check('failure fraction among observed records', ratio(5,90),float(Fraction(50,9)))
 check('unknown records retained separately', statuses.count('unknown'),10)
 expected = ['a','b','c','d']
 rows = [{'id':'a','result':'pass'},{'id':'b','result':'pass'},{'id':'b','result':'pass'},{'id':'c','result':'fail'}]
 r = reconcile(expected, rows)
 check('raw row count can equal population', len(rows), len(expected))
 check('duplicate cannot fill missing observation', r['missing'],['d'])
 check('identical repeated evidence counted', r['duplicate_copies'],1)
 check('conflicting result isolated', reconcile(expected,rows+[{'id':'b','result':'fail'}])['conflicts'],['b'])
 check('unexpected evidence identity exposed',reconcile(expected,rows+[{'id':'x','result':'pass'}])['unexpected'],['x'])
 check('stable definition first period defect percent', ratio(10,1000),1)
 check('stable definition second period defect percent', ratio(15,3000),0.5)
 check('absolute exception count still rises',15-10,5)
 check('combined pass ratio uses combined counts',ratio(90+50,100+1000),float(Fraction(1400,110)))
 check('unweighted average differs', (90+5)/2,47.5)
 check('stale observations cannot prove current period', 60<=15,False)
 events=[dict(id='e1',kind='loss',amount=12000),dict(id='e2',kind='loss',amount=3000),dict(id='e1',kind='loss',amount=12000),dict(id='n1',kind='near-miss',amount=50000)]
 recovery=[dict(id='r1',event='e1',amount=4000,status='received'),dict(id='r2',event='e1',amount=2000,status='pending')]
 result=loss_summary(events,recovery)
 check('distinct event records include near miss',result['events'],3)
 check('gross ignores duplicate and potential near miss',result['gross'],15000)
 check('received recovery excludes pending claim',result['received'],4000)
 check('net observed amount retains gross lineage',result['net'],11000)
 for name,ev,rec,message in [
 ('conflicting event versions fail',events+[dict(id='e1',kind='loss',amount=13000)],recovery,'conflicting event version'),
 ('duplicate recovery identifiers fail',events,recovery+[recovery[0]],'duplicate recovery identifier'),
 ('unmatched recovery fails',events,[dict(id='r3',event='missing',amount=100,status='received')],'unmatched recovery')]:
 try:loss_summary(ev,rec);value='accepted'
 except ValueError as e:value=str(e)
 check(name,value,message)
 gate=dict(implemented=True,effective=False,owner_acceptance=True)
 check('implementation alone cannot close action',all(gate.values),False)
 check('all fictional closure criteria demonstrated',all({**gate,'effective':True}.values),True)
 check('expired exception at exact boundary',10<10,False)
 return dict(runtime=platform.python_version,scope='Synthetic evidence, ratios and event ledger only; not statistical assurance, an actual RCSA, accounting policy, capital calculation or regulatory submission.',passed=len(checks),checks=checks,runnerSha256=hashlib.sha256(Path(__file__).read_bytes).hexdigest)

if __name__ == '__main__':
 p=argparse.ArgumentParser;p.add_argument('--output',required=True)
 args=p.parse_args;Path(args.output).write_text(json.dumps(run,indent=2)+'\n')
IN PRACTICE

Example: twenty normal changes passed, but no emergency change was selected. Retain the results and make the gap explicit before weekend acceptance.

Common pitfalls

Confusing sample passes with statistical confidence; counting copies as coverage; updating dates without observation; closing effectiveness on deployment.

Related topics: Risk assessment and control effectiveness · Events, near misses, and response · Indicators, reporting, and acceptance

Take this idea with you

The conclusion should be reconstructible from the objective, population, and observations, including what remains unproven.

Create account

Reference: Assessing Security and Privacy Controls in Information Systems and Organizations · BigSavant operational risk professional assessment2026.10