Define the investigation question
At 18:42 during a fictional fund close, a shell alert appears on a reconciliation replica. A maintenance window exists, but the alert has not yet been matched to the authorized intervention. APS must sustain processing while security clarifies the activity. Start with observable questions: which instance ran the process, which identity initiated the operation, which files changed and which period has available data? Record hypotheses separately from collected facts. An approved ticket does not authorize every action on the same date. Compare actor, scope, time and expected action. A denied request may be a misconfigured diagnostic attempt or part of a sequence requiring escalation. Keep alternatives open until evidence supports a conclusion. Establish who decides containment, who validates continuity and when the business owner must become involved. This lesson connects API requests, container state and observation quality. Its exercise performs real operations in a disposable cluster using synthetic values. The business scenario, deadlines and decisions are original and do not represent internal BNP Paribas procedures. The objective is a defensible conclusion for the next operational decision, explaining observations, inferences and unresolved gaps. A useful handover should let another responder retrace that reasoning without relying on the first analyst’s memory.
Build demonstrable audit coverage
Start from the operation you need to observe and inspect effective policy. Selecting pods does not automatically cover pods/log. An analyst investigating who retrieved logs must confirm the subresource and audit backend; Events used for scheduling diagnosis provide different evidence. If an earlier rule already matches a request, a later detailed rule may never apply. Review ordering with concrete requests. The exercise keeps Secrets at Metadata level and synthetic ConfigMaps at RequestResponse. This allows comparison of a change request with its result without copying credential values into published evidence. In real systems, assess list responses too: a bodyless request can produce a collection containing sensitive data. Short retention reduces exposure duration but does not remove unnecessary copying. Record the union of global and per-rule omitStages before explaining a missing stage. A watch can produce different stages sharing an auditID; do not automatically count every line as a separate action. Check rotation, collection and retrieval of a known event across the complete path. A populated local file does not prove the collector retained rotated files or that an analyst can retrieve them for the incident interval. Assign an owner to that retrieval check so successful collection remains observable after configuration changes.
Correlate requests without inventing identity
The laboratory submits a request through impersonation permitted to its training-cluster administrator. Authorization of the impersonated operation is denied. The event retains user and impersonatedUser, separating the authenticated client from the identity used for the access decision. It does not turn a ServiceAccount into a person. Human attribution would require additional session and identity-system sources with explicit trust and limitations. User-Agent helps search for patterns but is client-supplied information. A kubectl string does not attest the executable or device. Likewise, sourceIPs requires knowledge of proxies and header origins; a forwarded address may not carry the trust of the last observed hop. Preserve the original record before normalizing search fields and document where enriched information came from. Do not force every request into an object schema containing metadata. JSON Patch can carry an operation array. Distinguish requestObject, responseObject and the event’s own annotations. When constructing a timeline, separate request time, stage time and collector ingestion. If delay or clock skew exists, state that uncertainty instead of inferring causality solely from dashboard arrival order. The sequence must remain understandable when another team reviews the data. Record which correlations are exact identifier matches and which depend on an approximate time window.
Distinguish image, filesystem and instance
A digest identifies image content but does not describe all writable state in an execution. The exercise changes /tmp/dr-state and observes that the Pod reference remains unchanged. It then compares a container using readOnlyRootFilesystem:true: writing under /tmp returns EROFS while /scratch, mounted from an emptyDir, remains writable. Those results identify concrete protection boundaries rather than a configuration contradiction. When an application needs writes, locate the path and establish its lifecycle. Distinguish disposable cache, approved configuration and persistent data. chmod does not convert a read-only mount into a writable one. Give a volume only necessary permissions and scope, with ownership of cleanup, capacity and recovery. Also check whether the application can load code or configuration from that area. To investigate a change, retain file, path, mount, Pod UID, container identity and time. A Pod name can reappear after recreation; the laboratory demonstrates a new UID. PIDs also need process-namespace context. A hash computed after an alert identifies collected bytes without establishing that they are clean. Compare against an approved origin and retain data needed for analysis under the authorized collection procedure. Describe any decision to release the instance and the evidence that decision may make unavailable.
Design detection with explicit scope
Before writing a Falco rule, identify the event source. Kubernetes request fields belong to the k8s_audit plugin; a syscall rule does not acquire those fields simply because it includes ka.verb. Sensor installation, plugin configuration and delivery need their own validation. This lesson’s exercise does not install Falco: rules are analyzed from primary documentation without claiming execution in the detection engine. Write the intent and examples before the condition. For A and (B or C), test an input satisfying only C while failing A to detect an overly broad grouping. An exception for rotate-agent under /var/app/checkpoints should retain the relationship between process and destination. Excluding only the process name can also silence writes outside the authorized path. Test an allowed action and a similar action that must still alert. A missing field is not necessarily the string <NA> confirm extraction semantics for the plugin version and use exists when the intent is presence. Priority communicates severity without demonstrating evaluation order or containment. If output feeds automation, separately validate its consumer, authorization, condition and result. The report should say which steps were actually exercised and which remain designs. Keep the rule version and test inputs with the review so another analyst can repeat the reasoning.
Measure confidence in observation
An alert-free dashboard can coincide with capture failure. During a processing spike, record growth in scap.n_drops and investigate load, buffers, event selection and rules. Removing rules can reduce work while also removing coverage. Compare the effect with the detection requirement instead of accepting improvement in the technical indicator alone. If a cumulative counter resets with the process, do not interpret a negative difference as recovery of dropped events. Separate lifecycles and calculate increases within each. Retain the restart time and the interval lacking sufficient observation. Also distinguish capture loss, output-queue loss and collector delay. Each boundary may need a different test; a Running Pod does not prove the message reached a searchable destination. Apply the same care to API auditing. A pods/exec event at Metadata level does not contain a complete stdout transcript. In the exercise, a marker is sent over stdin, printed to the client and searched for in the audit log, where it is absent. This demonstrates a limitation of that channel and configuration. It does not establish non-execution or replace authorized collection of processes, files or other telemetry. Include those limitations in the handover so the next team does not treat silence as proof.
Run the exercise and interpret results
Use a dedicated kind cluster with auditing configured, Docker, Python 3 and kubectl. Laboratory material includes audit-policy.json and setup instructions. Import a Python image beforehand and record the digest for the platform actually available. The script uses imagePullPolicy:Never and refuses contexts lacking the kind-dr-cks-runtime- prefix. Do not apply it to a shared environment: it creates ConfigMaps, synthetic Secrets and Pods in a namespace that must be absent. Save the code as run.py and execute python3 run.py --kubeconfig PATH --context CONTEXT --node CONTROL-PLANE-NODE --image REFERENCE@sha256:DIGEST --output NEW-RESULT.json. Check arguments against the laboratory README. The script refuses to overwrite results, removes the namespace at the end and retains execution events. Cluster removal is a separate step after saving necessary evidence. Trials use Kubernetes server 1.37.0 and client 1.37.1; the official CKS page states runtime 1.35. That difference is recorded and is not presented as validation of the exam runtime. Interpret the 14 observations within their scope: denied authorization, files, mounting, UID and audit events. This execution includes no Falco sensor, syscall capture, network containment, memory forensics or full practical mock. Repeat with a new output path and compare outcomes rather than expecting identical timestamps or object identifiers.
Decide containment and hand over investigation
In the second case, a reconciliation-volume file changed and the sensor dropped events during the same interval. An approved digest and read-only root filesystem do not resolve whether that change was legitimate. A validated alternative replica exists and the plan authorizes targeted containment. Coordinate context preservation with the approved measure, confirm continuity and record incomplete observation. The laboratory does not perform that containment; the case assesses the decision. Hand over a concise timeline, identifiers, sources, collected-file hashes and actions taken. Separate facts, hypotheses, decisions and next steps. Record who authorizes a change, who confirms capacity and when the measure must be reviewed. Avoid restarting automatically before collection merely to regain green status. If continuity is at risk, apply approved criteria and document the effect on evidence not yet collected. Summary: establish coverage before interpreting absence, retain distinct identities, correlate instances and assess filesystems and volumes separately. Detection rules need positive and negative tests; metrics need temporal context. Connect this lesson with RBAC, admission, Linux hardening, supply chain and incident management. CKS preparation still requires broader practical training and specialist review beyond this internal decision assessment. Ask the receiving team to explain both an observed result and one unresolved uncertainty before closing the handover.
#!/usr/bin/env python3
"""Real API-audit and filesystem exercise for a dedicated kind training cluster.
Requires the supplied audit policy and an already imported Python image digest.
Only synthetic ConfigMaps/Secrets and owned Pods are created. No Falco sensor,
network containment, forensic memory capture or production identity is used.
"""
import argparse, base64, datetime, hashlib, json, pathlib, subprocess, time, uuid
p = argparse.ArgumentParser(description=__doc__)
p.add_argument('--kubeconfig', required=True)
p.add_argument('--context', required=True)
p.add_argument('--node', required=True)
p.add_argument('--image', required=True)
p.add_argument('--output', required=True)
p.add_argument('--kubectl', default='kubectl', help='Compatible kubectl executable')
a = p.parse_args
if not a.context.startswith('kind-dr-cks-runtime-'):
raise SystemExit('Use a dedicated kind-dr-cks-runtime-* context.')
if a.node!= a.context.removeprefix('kind-') + '-control-plane':
raise SystemExit('Node must belong to the named dedicated cluster.')
if '@sha256:' not in a.image:
raise SystemExit('Provide an already imported image by digest; this lab never pulls.')
out = pathlib.Path(a.output).resolve
if out.exists:
raise SystemExit('Output already exists; choose a new report path.')
ns = 'dr-cks-runtime-lab'
prefix = 'trial-' + uuid.uuid4.hex[:8]
observations = []
def run(cmd, body=None, ok=True, timeout=40):
r = subprocess.run(cmd, input=body, text=True, capture_output=True, timeout=timeout)
if ok and r.returncode:
raise RuntimeError(f'{cmd[0]} failed: {r.stderr[-1500:]}')
return r
def k(*args, body=None, ok=True, timeout=40):
return run([a.kubectl, '--kubeconfig', a.kubeconfig, '--context', a.context, *args], body, ok, timeout)
def create(obj):
return json.loads(k('create', '-f', '-', '-o', 'json', body=json.dumps(obj)).stdout)
def observe(name, passed, **values):
observations.append(dict(name=name, passed=bool(passed), observed=values))
if not passed:
raise AssertionError(name + ': ' + str(values))
def pod(name, readonly):
return dict(apiVersion='v1', kind='Pod', metadata=dict(name=name, namespace=ns),
spec=dict(automountServiceAccountToken=False, restartPolicy='Never',
securityContext=dict(runAsNonRoot=True, runAsUser=10001, runAsGroup=10001,
fsGroup=10001, seccompProfile=dict(type='RuntimeDefault')),
containers=[dict(name='worker', image=a.image, imagePullPolicy='Never',
command=['python3', '-c', 'import time; time.sleep(900)'],
securityContext=dict(allowPrivilegeEscalation=False,
readOnlyRootFilesystem=readonly, capabilities=dict(drop=['ALL'])),
volumeMounts=[dict(name='scratch', mountPath='/scratch')])],
volumes=[dict(name='scratch', emptyDir={})]))
def execute(name, script):
return k('-n', ns, 'exec', name, '-c', 'worker', '--', 'python3', '-c', script)
report = dict(startedAt=datetime.datetime.now(datetime.timezone.utc).isoformat,
scriptSha256=hashlib.sha256(pathlib.Path(__file__).read_bytes).hexdigest,
version=json.loads(k('version', '-o', 'json').stdout), context=a.context,
imageReference=a.image, namespace=ns, trial=prefix, observations=observations,
examVersion='1.35', fullPracticalMock=False, independentVerification=False,
scope='Actual isolated API audit events, authorization denial and container filesystem writes. No Falco engine, syscall capture, network isolation, memory forensics or production incident.')
client_minor = int(report['version']['clientVersion']['minor'].rstrip('+'))
server_minor = int(report['version']['serverVersion']['minor'].rstrip('+'))
if abs(client_minor - server_minor) > 1:
raise SystemExit('Use kubectl within one minor version of the API server; set --kubectl explicitly.')
existing = k('get', 'namespace', ns, ok=False)
if existing.returncode == 0 or 'NotFound' not in existing.stderr:
raise SystemExit('Namespace exists or absence is unconfirmed; refusing to reuse it.')
owned = False
try:
create(dict(apiVersion='v1', kind='Namespace', metadata=dict(name=ns)))
owned = True
cm = prefix + '-settings'
secret = prefix + '-credential'
marker = 'synthetic-only-' + uuid.uuid4.hex
config = create(dict(apiVersion='v1', kind='ConfigMap', metadata=dict(name=cm, namespace=ns), data=dict(mode='approved')))
create(dict(apiVersion='v1', kind='Secret', metadata=dict(name=secret, namespace=ns), stringData=dict(token=marker)))
k('-n', ns, 'get', 'secret', secret, '-o', 'json')
k('-n', ns, 'patch', 'configmap', cm, '--type=json', '-p', '[{"op":"replace","path":"/data/mode","value":"changed"}]')
impersonation = 'system:serviceaccount:' + ns + ':unprivileged-reader'
denied = k('-n', ns, 'get', 'configmap', cm, '--as', impersonation, ok=False)
observe('unauthorized-read-denied', denied.returncode!= 0 and 'Forbidden' in denied.stderr,
clientExitCode=denied.returncode, impersonatedIdentity=impersonation)
mutable = prefix + '-mutable'
immutable = prefix + '-readonly'
initial = create(pod(mutable, False))
create(pod(immutable, True))
k('-n', ns, 'wait', '--for=condition=Ready', 'pod/' + mutable, 'pod/' + immutable, '--timeout=90s', timeout=100)
write = "import pathlib,hashlib,json; p=pathlib.Path('/tmp/dr-state'); p.write_text('approved'); before=hashlib.sha256(p.read_bytes).hexdigest; p.write_text('changed'); after=hashlib.sha256(p.read_bytes).hexdigest; print(json.dumps({'before':before,'after':after,'changed':before!=after}))"
change = json.loads(execute(mutable, write).stdout)
observe('mutable-root-file-changed', change['changed'], **change)
checked = json.loads(k('-n', ns, 'get', 'pod', mutable, '-o', 'json').stdout)
observe('image-reference-unchanged-after-write', checked['spec']['containers'][0]['image'] == a.image,
declaredImage=checked['spec']['containers'][0]['image'], uid=checked['metadata']['uid'],
imageID=checked['status']['containerStatuses'][0]['imageID'])
readonly = "import pathlib,json,errno\ntry:\n pathlib.Path('/tmp/dr-state').write_text('changed'); result={'denied':False}\nexcept OSError as e:\n result={'denied':e.errno==errno.EROFS,'errno':e.errno}\nprint(json.dumps(result))"
result = json.loads(execute(immutable, readonly).stdout)
observe('readonly-root-write-denied', result['denied'], **result)
scratch = "import pathlib,json; p=pathlib.Path('/scratch/dr-state'); p.write_text('synthetic'); print(json.dumps({'writable':p.read_text=='synthetic'}))"
result = json.loads(execute(immutable, scratch).stdout)
observe('explicit-volume-still-writable', result['writable'], **result)
# Prove that the API records an exec request, not a transcript of its stdout.
stdout_marker = 'stdout-only-' + uuid.uuid4.hex
result = k('-n', ns, 'exec', '-i', mutable, '-c', 'worker', '--', 'python3', '-c',
'import sys; print(sys.stdin.read)', body=stdout_marker)
observe('exec-returned-synthetic-output', stdout_marker in result.stdout, outputMatched=True)
# Watch uses an API timeout and generates stages that can share an auditID.
uri = '/api/v1/namespaces/' + ns + '/configmaps?watch=true&timeoutSeconds=2'
k('get', '--raw', uri, timeout=10)
old_uid = checked['metadata']['uid']
k('-n', ns, 'delete', 'pod', mutable, '--wait=true', '--timeout=45s', timeout=50)
recreated = create(pod(mutable, False))
observe('recreated-name-has-new-uid', recreated['metadata']['uid']!= old_uid,
objectName=mutable, previousUID=old_uid, currentUID=recreated['metadata']['uid'])
# Read only the owned control-plane log, then select this trial namespace.
events = []
for _ in range(30):
raw = run(['docker', 'exec', a.node, 'cat', '/var/log/dr-audit/audit.log']).stdout
candidates = [json.loads(line) for line in raw.splitlines if line.strip]
events = [e for e in candidates if e.get('objectRef', {}).get('namespace') == ns
and e.get('requestReceivedTimestamp', '') >= report['startedAt'].replace('+00:00', 'Z')]
if any(e.get('verb') == 'watch' and e.get('stage') == 'ResponseComplete' for e in events):
break
time.sleep(0.2)
def select(resource, verb=None, name=None):
return [e for e in events if e.get('objectRef', {}).get('resource') == resource
and (verb is None or e.get('verb') == verb)
and (name is None or e.get('objectRef', {}).get('name') == name)]
created = select('configmaps', 'create', cm)
observe('configmap-request-response-recorded', any(e.get('level') == 'RequestResponse'
and e.get('requestObject', {}).get('data', {}).get('mode') == 'approved'
and e.get('responseObject', {}).get('metadata', {}).get('uid') == config['metadata']['uid']
and e.get('responseStatus', {}).get('code') == 201 for e in created), events=len(created))
patches = select('configmaps', 'patch', cm)
observe('json-patch-array-recorded', any(isinstance(e.get('requestObject'), list)
and e.get('responseObject', {}).get('data', {}).get('mode') == 'changed'
and e.get('responseStatus', {}).get('code') == 200 for e in patches), events=len(patches))
secret_events = select('secrets', name=secret)
observe('secret-bodies-omitted', len(secret_events) >= 2 and marker not in raw
and base64.b64encode(marker.encode).decode not in raw and all(e['level'] == 'Metadata'
and 'requestObject' not in e and 'responseObject' not in e for e in secret_events),
events=len(secret_events), secretMarkerAbsent=marker not in raw,
encodedSecretMarkerAbsent=base64.b64encode(marker.encode).decode not in raw)
denied_events = [e for e in select('configmaps', 'get', cm) if e.get('responseStatus', {}).get('code') == 403]
observe('impersonation-attribution-retained', any(e.get('impersonatedUser', {}).get('username') == impersonation
and e.get('user', {}).get('username')!= impersonation for e in denied_events),
attribution=[dict(user=e.get('user', {}).get('username'),
impersonatedUser=e.get('impersonatedUser', {}).get('username'),
auditID=e.get('auditID'), code=e.get('responseStatus', {}).get('code')) for e in denied_events])
watches = select('configmaps', 'watch')
grouped = {}
for e in watches:
grouped.setdefault(e['auditID'], []).append(e['stage'])
observe('watch-stages-share-request-id', any('ResponseStarted' in stages and 'ResponseComplete' in stages
for stages in grouped.values), requestStages=grouped)
observe('requestreceived-omitted-by-policy', bool(events) and all(e['stage']!= 'RequestReceived' for e in events),
retainedEvents=len(events))
exec_events = [e for e in select('pods') if e.get('objectRef', {}).get('subresource') == 'exec']
observe('exec-metadata-without-stdout-transcript', bool(exec_events) and stdout_marker not in raw,
execAuditEvents=len(exec_events), stdoutMarkerAbsent=stdout_marker not in raw)
report['events'] = events
report['auditExportSha256'] = hashlib.sha256(json.dumps(events, sort_keys=True).encode).hexdigest
report['passed'] = all(o['passed'] for o in observations)
finally:
if owned:
k('delete', 'namespace', ns, '--wait=true', '--timeout=90s', timeout=100)
absent = k('get', 'namespace', ns, ok=False)
report['cleanup'] = dict(namespaceRemoved=absent.returncode!= 0 and 'NotFound' in absent.stderr,
realCredentialsUsed=False, clusterDeletionSeparate=True)
report['finishedAt'] = datetime.datetime.now(datetime.timezone.utc).isoformat
out.write_text(json.dumps(report, indent=2) + '\n')
print(json.dumps(dict(passed=report['passed'], observations=len(observations), report=str(out))))
A volume file changes despite an approved digest; the team preserves context, acknowledges dropped events and decides targeted containment.
Common pitfalls
Confusing images with volume state, names with UIDs, User-Agent with identity or absent alerts with complete coverage.
Related topics: RBAC and identities · Hardening and volumes · Incident response
A useful investigation preserves context, explains confidence in each source and supports the next operational decision.
Reference: CKS certification and domains · Kubernetes v1.35; current six-domain CKS outline