← CKS: Kubernetes security in production
23 / 23 · 120 MIN

Detect, contain and recover with evidence

Investigate detection gaps, validate rules and distinguish containment from service recovery.

Bound the incident and confidence in signals

During an incident, missing alerts have several possible explanations. Activity may not have occurred, collection may have failed, a rule may have been disabled or delivery may be delayed. Before concluding, identify the path from activity to decision: API server, audit policy, transport, plugin, rules, output and destination. For each boundary, record what was observed and what still needs confirmation. A healthy process does not establish that every step works. In a fictional banking case, APS supports fund closing while security investigates a Falco rule change. The business owner needs to know whether processing can continue and what risk the gap creates. Prepare a timeline containing the last known configuration, the change, the first unexpected observation and available retention windows. Do not attribute behavior to a person merely because an account appears in a record. Authenticated identity is a clue to relate to access, context and other evidence. Prepare concrete meeting questions: which operations should generate a signal, which were seen at source and which reached the destination? Which exception was authorized, for which account, target and interval? What continuity decision applies until the next checkpoint? This is an educational case and does not describe internal BNP Paribas procedures. It develops coordination between operations, security and business without turning incomplete information into an assurance. Assign an owner to each missing fact so the next shift can continue the investigation from evidence.

Review the rules actually applied

A correct rule in a file may not participate in the intended evaluation. First identify Falco version, loaded plugins, effective files and their order. An override that appends a condition depends on an earlier definition. Replace deserves particular care: replacing the condition may lose namespace, resource or verb limits established by its previous author. Review the entire resulting expression, including parentheses, rather than approving only the changed line. With rule_matching: first, a general rule can match before a specific rule. Raising the latter’s priority does not change ordering. All permits further matches but can increase volume and make several rules describe the same event. Distinguish events, alerts and incidents in reporting. The global priority threshold has a separate effect: priority: error excludes a WARNING rule. Log_level: debug increases operational detail without expanding the rule set selected by that threshold. A maintenance exception should represent the authorized combination. If ops-a may act only in the test namespace, excluding the whole account can suppress activity outside that scope. Prepare examples differing in one field: the same account in another namespace, another account against the authorized target and another verb on the same resource. These controls expose boundaries that a single positive example cannot establish. Also record command-line enable and disable options, whose order contributes to effective behavior. Retain both the reviewed files and the invocation used for the decision.

Separate validation, replay and real operation

Validating syntax and rule dependencies is necessary but does not establish that input produces the expected result. A dry run does not process events to demonstrate matching. In addition, skip-if-unknown-filter: true may allow a rule to load without executing when it contains an unknown field. Confirm fields from the actually installed plugin, observe an input that should match and include inputs that should be excluded. No alert for a negative case does not distinguish a correct rule from an inactive rule. For the proposed exercise, use only original synthetic audit.k8s.io/v1 events. Give each request a distinct auditID and include stage, identity, verb, target and outcome where applicable. The k8saudit plugin can read a plain local path to end of file; file:// follows the file continuously and must not be confused with that finite input. Retain plugin version, configuration, rules and an input hash to repeat the comparison. Plan tests for rule order, bounded exception, override and minimum severity. Compare observed identifiers with expectations. The laboratory ran Falco 0.45.0 with k8saudit 0.18.0 in two final runs of 41 checks each, using exactly the displayed code. It confirmed matches, exclusions and expected failures; an initial run showed that the version check needed to parse JSON. Continue to assess the actual API-server-to-destination transport, syscall capture and cluster containment separately. Keep those boundaries explicit in the readiness report. An observation should support a specific claim about the path exercised, with missing paths left visible for subsequent work. Image and plugin are pinned by digest and checksum. The exercise runs without networking, as an unprivileged user, with a read-only filesystem and only synthetic files mounted. It removes its containers and temporary files afterwards. The library and image must be available before running the command described in the code.

Investigate loss and delay between components

The collection chain contains different queues and buffers. In batch mode, API-server auditing exports events asynchronously. A collector pause can grow a queue before any Falco rule sees the data. In a sizing example, 120 requests per second with two audited stages generate 1200 events in five seconds. That calculation ignores bursts, other events, sizes and queue recovery; it identifies requirements rather than certifying capacity. In Falco, the queue between engine and output channels differs from buffering inside a channel. Bounded capacity can drop events when exhausted; an unlimited queue also creates memory risk. Disabling buffering may reduce stdout flush delay at operational cost but does not recover events already lost upstream. Investigate producer metrics, delivery health, sensor resources and destination persistence. The apiserver_audit_error_total counter helps observe export losses without replacing investigation of the affected interval. From Falco 0.45, handling invalid bytes and control characters also warrants parser review. Conditions evaluate original input while output may represent characters differently. Avoid reconstructing original bytes from normalized text when investigation requires fidelity. Use fictional data to exercise parsing and preserve the input file. Document the difference between delay, loss and representation change because each requires a different response. When briefing the incident lead, describe the affected interval and the evidence still available rather than reporting every missing alert as a rule defect.

Correlate a request with its outcome

An audited request can generate events at several stages. RequestReceived shows receipt rather than successful completion. If a later outcome has code 403, the initial record should not be presented as evidence of authorized access. Preserve the relationship between auditID and stages. A collector retaining only the first event per identifier may discard exactly the outcome information needed for investigation. Deduplication must preserve meaningful differences rather than merely reduce lines. In a control plane with several API servers, a load-balanced query may miss the instance with outdated policy. Confirm policy adoption and record origin per instance. Also retain cluster context when aggregating data. A successful test on one path does not automatically cover another with different configuration. If the destination has a gap, consult remaining sources and state reconstruction limits; do not invent successful outcomes for requests lacking confirmation. For guided application, prepare a table containing request, stage, target, observed outcome, origin and uncertainty. Use three fictional situations: receipt followed by 403, successful completion and receipt alone during export failure. Explain what each row establishes and what it cannot establish. Carry this distinction into handover: “observed attempt” and “completed operation” imply different decisions. Relate the interpretation to access control, evidence retention and service-impact analysis. Preserve the original record alongside any normalized investigation view so another analyst can reassess the conclusion.

Demonstrate containment effects

An action accepted by the API may have a narrower effect than the team expects. Rollout pause prevents rollout progression but does not stop processes in existing replicas. Force-deleting a Pod also does not wait for kubelet confirmation of termination. On an unreachable node, the object may disappear while execution continues. Before recording “isolated,” define the intended boundary and the observation that will establish it. For NetworkPolicy, consider every applicable permission. A quarantine policy with no egress rules does not cancel permission from another policy selecting the same Pod. Blocked new connections also do not establish termination of an earlier session: existing-connection behavior depends on implementation. HostNetwork behavior needs its own plugin assessment. Object acceptance alone does not supply the instant when the dataplane adopted a change. In a decision exercise, the reconciliation Pod cannot open new connections but retains a suspect session to an external destination. Response must address that residual path while preserving access needed for investigation. Identify authority for a service-impacting action, expected consequence and subsequent observation. Avoid alternating policy changes without a clear hypothesis because that makes sequence reconstruction harder. These limits are studied from documentation; local Falco event replay neither executes nor demonstrates cluster network containment. Write the test boundary in the evidence record so a successful detector exercise is not reused as proof of an unrelated control.

Recover without losing data or evidence

An approved image does not automatically make mounted data safe. If an incident changed configuration on a persistent volume, replacing the Pod while reusing that volume may restore the same behavior. Identify each configuration’s origin, affected business data and an acceptable recovery state. Preserve what investigation and reconciliation need before destructive action. Restarting does not reverse completed payments. EmptyDir has a different lifetime: it preserves data through container crashes in the same Pod, but removing the Pod loses that content. If it holds the only useful diagnostic files and an authorized window remains, collect them before replacement. Record their Pod, container and collection-time context. Urgent containment may limit what can be preserved; state the decision and lost information rather than promising later reconstruction. For a distroless image, an ephemeral container can provide tools without embedding them in the application image. Assess image, access, permissions and process visibility. The mechanism does not provide automatic restart, and its entry cannot be changed or removed after addition. Record the intervention even when the diagnostic process exits. As an exercise, distinguish three decisions: collect evidence, stop suspect execution and restore consistent processing. Each has its own criteria and may require coordination with different teams. A green health check helps evaluate availability but does not resolve the integrity of mounted configuration or the outcome of earlier business transactions.

Prepare handover and recovery criteria

Handover should let the next team continue from facts. Provide rule and plugin versions, investigated interval, relevant requests, known losses, observed containment and paths still awaiting tests. Separate completed actions from decisions authorized but not yet executed. For each uncertainty, identify possible impact, owner and next checkpoint. Include the change needed to restore earlier configuration when a detection change fails. In the payments case, a healthy process and approved image do not authorize indiscriminate batch replay. Confirm transaction state with the responsible team, identify repeatable items and establish duplicate-prevention controls. If evidence cannot support a decision before the window, present an assessed fallback or a batch hold with communicated impact. This applies recovery reasoning to a professional context rather than prescribing one institution’s operating rules. Complete the exercise by delivering a five-column matrix: claim, evidence, limitation, decision and owner. Example claims include “the rule matches again,” “the destination received the alert,” “the suspect session ended” and “the batch is reconciled.” Require different evidence for each. The lesson’s conclusion is straightforward: restoring the sensor, containing execution and recovering service are related tasks, but each conclusion needs its own observation. Connect them to auditing, immutability, incident management, change and continuity. Retain unfinished checks explicitly so a later report can distinguish new evidence from a repeated assumption.

#!/usr/bin/env python3
"""Original synthetic Kubernetes audit replay. No cluster, host sensors or real credentials.
Requires Docker, the pinned Falco image and the checksum-verified ARM64 plugin.
Usage: python3 run.py --plugin /path/libk8saudit.so --report /path/evidence.json
"""
import argparse,datetime,hashlib,json,pathlib,shutil,subprocess,tempfile,uuid
IMAGE='falcosecurity/falco@sha256:788f1129c542171813083d4afc61b16730a47dde8c23d9c39370acef996349b6'
PLUGIN_SHA='a01cf425b9c345b6908196fe828d363531a3f1261e6346e4e8f4c25a281eccd3'
BASE='ka.target.resource=secrets and ka.verb=get and ka.stage=ResponseComplete'
def sha(data):return hashlib.sha256(data).hexdigest
def event(n,user='ops-a',namespace='funds',resource='secrets',verb='get',stage='ResponseComplete',code=200):
 return dict(apiVersion='audit.k8s.io/v1',kind='Event',level='Metadata',auditID=f'dr-{n:02}',stage=stage,verb=verb,user={'username':user},objectRef={'resource':resource,'namespace':namespace,'name':'synthetic-only'},responseStatus={'code':code},requestReceivedTimestamp='2026-10-07T08:00:00.000000Z',stageTimestamp='2026-10-07T08:00:00.000001Z')
def rule(name,condition=BASE,**extra):
 return dict(rule=name,desc='Original BigSavant synthetic event exercise',condition=condition,output='DR request=%ka.auditid user=%ka.user.name namespace=%ka.target.namespace code=%ka.response.code',priority='WARNING',source='k8s_audit',**extra)
def main:
 ap=argparse.ArgumentParser;ap.add_argument('--plugin',type=pathlib.Path,required=True);ap.add_argument('--report',type=pathlib.Path,required=True);a=ap.parse_args
 assert sha(a.plugin.read_bytes)==PLUGIN_SHA,'Unexpected plugin bytes'
 assert not a.report.exists,'Report must be a new path'
 a.report.parent.mkdir(parents=True,exist_ok=True)
 work=pathlib.Path(tempfile.mkdtemp(prefix='dr-cks-falco-'));work.chmod(0o755)
 shutil.copyfile(a.plugin,work/'libk8saudit.so');(work/'libk8saudit.so').chmod(0o644)
 name='dr-cks-falco-'+uuid.uuid4.hex[:12];observations=[];runs=[]
 report=dict(startedAt=datetime.datetime.now(datetime.timezone.utc).isoformat,scriptSha256=sha(pathlib.Path(__file__).read_bytes),image=IMAGE,pluginSha256=PLUGIN_SHA,syntheticOnly=True,actualFalcoReplay=True,liveCluster=False,syscallCapture=False,productionWebhook=False,CNIContainment=False,fullPracticalMock=False,observations=observations,runs=runs)
 events=[event(1),event(2,namespace='test'),event(3,user='ops-b'),event(4,resource='configmaps'),event(5,verb='list'),event(6,stage='RequestReceived'),event(7,code=403),event(8,user='ops-a\nFORGED-LINE')]
 (work/'events.jsonl').write_text(''.join(json.dumps(x)+'\n' for x in events));report['inputSha256']=sha((work/'events.jsonl').read_bytes)
 def check(label,actual,expected):
 observations.append(dict(name=label,observed=actual,expected=expected,passed=actual==expected));assert actual==expected,(label,actual,expected)
 def run(label,rule_sets,extra=None,argv=None,expected=None,failure=False):
 filenames=[]
 for i,rules in enumerate(rule_sets):
 f=work/f'rules-{i}.yaml'f.write_text(json.dumps(rules,indent=2));filenames.append('/lab/'+f.name)
 config=dict(engine={'kind':'nodriver'},plugins_hostinfo=False,plugins=[dict(name='k8saudit',library_path='/lab/libk8saudit.so',init_config=json.dumps({'maxEventSize':262144,'useAsync':False}),open_params='/lab/events.jsonl')],load_plugins=['k8saudit'],rules_files=filenames,json_output=True,buffered_outputs=False,stdout_output={'enabled':True},syslog_output={'enabled':False},webserver={'enabled':False},rule_matching='first',priority='debug')
 if extra:config.update(extra)
 (work/'falco.yaml').write_text(json.dumps(config,indent=2))
 cmd=['docker','run','--rm','--name',name,'--network','none','--read-only','--cap-drop','ALL','--security-opt','no-new-privileges','--user','65534:65534','--memory','256m','--pids-limit','64','--tmpfs','/tmp:rw,noexec,nosuid,size=16m','-v',str(work)+':/lab:ro','--entrypoint','/usr/bin/falco',IMAGE,'-c','/lab/falco.yaml','--enable-source','k8s_audit',*(argv or [])]
 r=subprocess.run(cmd,capture_output=True,text=True,timeout=30)
 alerts=[]
 for line in r.stdout.splitlines:
 try:d=json.loads(line)
 except ValueError:continue
 if isinstance(d,dict) and d.get('source')=='k8s_audit' and 'rule' in d and 'output_fields' in d:alerts.append(d)
 facts=sorted([[x['rule'],x['output_fields']['ka.auditid']] for x in alerts])
 runs.append(dict(name=label,exitCode=r.returncode,config=config,rules=rule_sets,arguments=argv or [],stdout=r.stdout,stderr=r.stderr,alerts=alerts))
 check(label+'-exit',r.returncode!=0 if failure else r.returncode,True if failure else 0)
 if expected is not None:check(label+'-matches',facts,sorted(expected))
 return alerts,r
 try:
 generic=rule('DR General');specific=rule('DR Specific',BASE+' and ka.user.name=ops-a')
 ids=[1,2,3,7,8];pairs=lambda rule_name,ids:[[rule_name,f'dr-{n:02}'] for n in ids]
 run('first-match',[[generic,specific]],expected=pairs('DR General',ids))
 run('all-match',[[generic,specific]],extra={'rule_matching':'all'},expected=pairs('DR General',ids)+pairs('DR Specific',[1,2,7]))
 run('reordered-first',[[specific,generic]],expected=pairs('DR Specific',[1,2,7])+pairs('DR General',[3,8]))
 ex=rule('DR Exception',exceptions=[{'name':'maintenance','fields':['ka.user.name','ka.target.namespace'],'comps':['=','='],'values':[['ops-a','test']]}])
 run('tuple-exception',[[ex]],expected=pairs('DR Exception',[1,3,7,8]))
 override={'rule':'DR General','condition':'and ka.target.namespace=funds','override':{'condition':'append'}}
 run('append-after',[[generic],[override]],expected=pairs('DR General',[1,3,7,8]))
 _,bad=run('append-before',[[override],[generic]],expected=[],failure=True)
 check('append-before-error-identifies-rule','DR General' in bad.stderr+bad.stdout,True)
 replace={'rule':'DR General','condition':'ka.user.name=ops-a','override':{'condition':'replace'}}
 run('replace-loses-bounds',[[generic],[replace]],expected=pairs('DR General',[1,2,4,5,6,7]))
 run('disable-then-enable',[[generic]],argv=['-o','rules[].disable.rule=*','-o','rules[].enable.rule=DR General'],expected=pairs('DR General',ids))
 run('enable-then-disable',[[generic]],argv=['-o','rules[].enable.rule=DR General','-o','rules[].disable.rule=*'],expected=[])
 run('severity-excludes',[[generic]],extra={'priority':'error','log_level':'debug'},expected=[])
 run('severity-includes',[[generic]],extra={'priority':'warning','log_level':'error'},expected=pairs('DR General',ids))
 unknown=rule('DR Unknown','ka.nonexistent_dr_field=example')
 _,bad=run('unknown-strict',[[generic,unknown]],failure=True,expected=[])
 check('unknown-strict-error-identifies-field','ka.nonexistent_dr_field' in bad.stderr+bad.stdout,True)
 unknown['skip-if-unknown-filter']=True
 run('unknown-skipped',[[unknown,generic]],expected=pairs('DR General',ids))
 run('dry-run',[[generic]],argv=['--dry-run'],expected=[])
 allowed=rule('DR Successful',BASE+' and ka.response.code="200"')
 run('response-success',[[allowed]],expected=pairs('DR Successful',[1,2,3,8]))
 alerts,_=run('json-control-character',[[generic]],expected=pairs('DR General',ids))
 changed=next(x for x in alerts if x['output_fields']['ka.auditid']=='dr-08')
 check('json-user-roundtrip',changed['output_fields']['ka.user.name'],'ops-a\nFORGED-LINE')
 check('single-alert-for-control-character',sum(x['output_fields']['ka.auditid']=='dr-08' for x in alerts),1)
 _,version=run('version',[[generic]],argv=['--version'])
 report['version']=json.loads(version.stdout);check('falco-version',report['version']['falco_version'],'0.45.0')
 _,plugin=run('plugin-info',[[generic]],argv=['--plugin-info','k8saudit']);report['pluginOutput']=plugin.stdout+plugin.stderr;check('plugin-version','0.18.0' in report['pluginOutput'],True)
 check('input-unchanged',sha((work/'events.jsonl').read_bytes),report['inputSha256'])
 report['passed']=True
 except BaseException as exc:
 report['passed']=False;report['error']=str(exc);raise
 finally:
 subprocess.run(['docker','rm','-f',name],capture_output=True,timeout=15)
 r=subprocess.run(['docker','ps','-a','--filter','name=^/'+name+'$','--format','{{.Names}}'],capture_output=True,text=True,timeout=15)
 shutil.rmtree(work)
 report['cleanup']=dict(ownedContainersRemoved=r.returncode==0 and not r.stdout.strip,temporaryMaterialRemoved=not work.exists,realCredentialsUsed=False,publicRegistryWrites=False)
 report['finishedAt']=datetime.datetime.now(datetime.timezone.utc).isoformat;a.report.write_text(json.dumps(report,ensure_ascii=False,indent=2)+'\n')
 print(json.dumps({'passed':report['passed'],'observations':len(observations),'report':str(a.report)}))
if __name__=='__main__':main
IN PRACTICE

The sensor is healthy, but an override broadened an exception and hid access outside the authorized namespace.

Common pitfalls

Confusing no alerts with no activity, rollout pause with containment and process health with data integrity.

Related topics: Auditing and observability · Runtime immutability · Incident response · Recovery and reconciliation

Take this idea with you

Every conclusion needs evidence from the boundary it claims to control.

Create account

Reference: CKS domains and exam details · Kubernetes v1.35; current six-domain CKS outline

Kubernetes® and CKS are trademarks or registered trademarks of The Linux Foundation. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by The Linux Foundation. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.