1. Recover service while preserving controls
A fictional APS team is preparing fund closing. During planned certificate rotation, new clients work against one replica and fail against another. Meanwhile, the previous application version can no longer create Pods under current policy. The technical manager needs to separate trust problems, configuration adoption and recovery capacity. These examples do not represent internal BNP Paribas procedures. Start with a simple matrix: consumer, destination, server generation, presented identity, TLS result and operation result. A certificate-verification error, an HTTP 403 and a Pod that never started belong to different stages. Collect evidence at the correct point and assign investigation to the team controlling that point. Do not substitute an aggregate green status for observation of the failing flow. Also define what must remain true during recovery: intended clients can operate, unauthorized identities remain denied, data are not duplicated and RUN can explain which version is serving. Planned rotation may permit an approved trust overlap; CA compromise requires a different containment analysis. Record that distinction before choosing fallback. Operational priority does not remove the security requirement, but it helps determine sequence and escalation time. The change record should name the person making each decision and the evidence they require. A successful retry against an unidentified replica is insufficient to close an incident involving inconsistent generations.
2. Prepare an admissible recovery version
Kubernetes recovery must pass controls in force when Pods are created. Retaining the previous ReplicaSet does not guarantee its template remains admissible. In the decision exercise, the previous release uses privileges now forbidden. The team should prepare a compatible image and template, validate dependencies and test creation before the change window. Inspect the rejected field. Pod Security controls also cover init and ephemeral containers; adding a privileged diagnostic tool may be denied even when the original application was compatible. Emergency access needs defined scope and approval. Exempting a RuntimeClass has another relevant effect: covered requests skip enforce, audit and warn behavior. Missing warnings in that situation are not a positive assessment. Restricted is also not a complete list of organizational requirements. For example, admission under that policy does not establish a read-only root filesystem. Check that requirement separately, including where the application needs to write. For the reconciliation case, add a functional checkpoint: what was the last confirmed movement and how can processing resume without duplication? The rollback plan should include compatible configuration, available dependencies, tests and a named decision owner. This lesson’s practice runs Restricted Pods but does not reproduce the admission and rollback scenarios described in this section. Treat those scenarios as preparation for environment-specific validation, with their own evidence still required.
3. Distinguish Secret delivery, adoption and recovery
Draw the material’s path: Secret object, volume projection, process read and loaded state. Each arrow is a stage to verify. In this exercise, the Python server reads its CA bundle when creating SSLContext and implements no reload. A later volume update can be correct while the server continues using previous trust. File hash and the result of a fresh connection answer different questions. For key and certificate pairs, prepare coherent material and validate loading before promotion. Updating two objects in one command does not create an application adoption transaction. A strategy with identified versions and controlled replacement lets you associate each consumer with its intended set. A missing required Secret should continue blocking startup that would otherwise be insecure; optional does not add authentication to the application. At-rest encryption recovery has its own dependencies. Before promoting a write key, relevant API servers need the ability to read that generation’s data. Rolling back to configuration that only knows the old key may fail when reading new objects. Beyond the current cluster, consider retained backups and the keys they depend on. Required evidence includes restoration capability, not only current writes. This lesson does not execute etcd encryption, KMS or backup restoration; it uses these situations to train recovery decisions. Assign ownership of key retention separately from the application release so neither team assumes the other has retained recovery material.
4. Preserve isolation boundaries in fallback
Imagine a service executing a partner’s code in an approved sandbox runtime. Its dedicated pool becomes unavailable. Removing runtimeClassName to start under the normal runtime changes an architectural security property. The decision should consider compatible alternative capacity or an explicit risk exception; it should not be hidden in a recovery patch. Capacity estimates should include RuntimeClass overhead as well as container requests. A pool exactly fitting requests may not fit the actual workload. Also confirm the installed handler and node eligibility. A fallback exercise should establish that approved configuration can actually execute and complete its flow within available capacity. Data and network isolation need separate checks. A retained PersistentVolume can outlive its namespace and still contain the previous tenant’s data. Reuse requires appropriate retention and sanitization handling. A NetworkPolicy accepted by the API does not demonstrate enforcement if the plugin does not implement it. When moving to fallback, test legitimate and forbidden traffic and identify the component enforcing control. The practice below uses ordinary networking in a kind cluster and installs neither a sandbox nor a NetworkPolicy engine. Application mTLS does not provide evidence for those mechanisms. Keep those exercises as additional practical work. Ask the receiving operations team to identify which isolation assumptions were validated and which remain unresolved before accepting recovery readiness.
5. Run Pod-to-Pod mTLS and separate authorization
The exercise uses a disposable cluster, short-lived synthetic certificates and a local Python image pinned by digest. The script requires a dedicated context and kubeconfig, confirms a local API server and creates a random namespace. It uses no real credentials. The official exam states Kubernetes 1.35; execution uses server 1.37.0 and kubectl 1.37.1, with Pod Security Restricted pinned to v1.35. These versions are not presented as equivalent in every behavior. The server listens on 8443, requires a client certificate and validates its chain against the server’s bundle. The client validates the server CA and expected name even while connecting to the Pod IP. The server certificate has a DNS SAN and serverAuth purpose; client certificates have URI SANs and clientAuth. They are locally issued for the exercise, without automatic binding to ServiceAccounts or a real SPIFFE issuer. After TLS, the application permits only the synthetic reconciler URI. The authorized client receives 200 and reporter receives 403 despite a valid certificate. With no certificate or an unknown issuer, the flow fails before an HTTP response. An incorrect expected name or incorrect server trust produces client verification failure. Before execution, predict these outcomes in a table. Use responses to distinguish transport, authentication and authorization while preserving the limitations of this simple program. Its exact source is included below so the behavior can be inspected and repeated.
6. Observe trust overlap and retirement
The first server generation accepts the old client CA. The new client is rejected. The script adds the new CA to the Secret and waits until the old server’s volume shows the updated hash. That process nevertheless continues rejecting the new client. Delivery occurred; the loaded SSLContext did not change. Measured projection delay belongs to this execution and does not define an SLA. The script then creates an overlap generation loading the updated bundle. Both old and new clients receive 200. Next, the Secret contains only the new CA. A third generation reads that set and rejects the old client while keeping the new one functional. Meanwhile, previous servers can still accept old clients: each retains its own context. The exercise removes those two generations and again tests acceptance of the new client and rejection of the old one. The 20 checks cover fresh TLS connections and short requests. They do not test OCSP, CRL, persistent-session termination or automatic service-mesh rotation. They do not measure production traffic or guarantee a lossless rollout. The experiment supports better acceptance criteria: identify all active verifiers, observe adoption and retire old generations in a controlled way. In real operations, add draining, observability, checkpoints and tests of paths actually serving requests. Record the server generation with every result so a successful test cannot be mistakenly attributed to an untested replica.
7. Apply the concepts to mesh policy
This lesson’s Istio questions use primary documentation; Istio is not installed in the lab. PeerAuthentication concerns incoming mTLS requirements. AuthorizationPolicy defines access decisions. An authenticated identity does not imply permission for every operation, just as reporter received 403 in the synthetic program. The analogy supports reasoning but does not establish that an Istio dataplane was configured or tested. When reviewing a change, consider its complete semantics. A matching DENY takes precedence over ALLOW; another ALLOW does not correct that denial. An ALLOW policy without rules differs from one containing an empty rule. In DENY, missing HTTP attributes can match TCP traffic, so protocol and scope require attention. The principal field uses peer-certificate identity rather than an arbitrary client-supplied header. Migration between sidecar and ambient also needs separate validation: current documentation does not support DISABLE in ambient. Do not copy an exception merely because the resource has the same name. Record installed version, mode, namespace, selector, target and expected outcome before testing. Prepare an authorized client and one that should be denied for each rule. If results differ from intent, preserve the control while identifying the responsible rule and scope. Promote the change only after tests explain effective access. Keep this documentation-based reasoning separate from the application laboratory’s measured results in the handover.
8. Close the change with evidence RUN can use
The technical manager’s deliverable is a supported, repeatable decision. For the closing case, present the consumer/generation matrix, approved overlap window, positive and negative tests, functional resumption point and owners. If the old CA were compromised, planned overlap could not automatically be reused as containment. That changed context requires a new risk decision. After the lab, confirm removal of the namespace, key files, disposable cluster and dedicated kubeconfig. The report retains versions, image digest, script hash, outcomes and limitations without storing private keys. Compare the two final executions and explain timing differences without converting an observation into a guaranteed deadline. Script success should be traceable through checks rather than depend on an isolated PASS message. Summarize learning as a sequence: identify the requirement, locate the verifier, observe delivered configuration, confirm adoption, test access and recover in a controlled way. Related topics are admission, Secrets, identity, authorization, isolation, capacity and recovery. Questions train original decisions; they are not real exam items. The internal assessment awards no official certification. Additional mesh, CNI and sandbox practice, full mocks and independent specialist review remain necessary for mature preparation. RUN should receive the evidence needed to repeat bounded checks, plus a clear list of mechanisms that this exercise did not demonstrate.
"""Original disposable-cluster lab: application mTLS, authorization and trust rotation."""
import argparse,datetime,hashlib,json,os,shutil,subprocess,tempfile,time,uuid
from pathlib import Path
p=argparse.ArgumentParser;p.add_argument('--kubectl',required=True);p.add_argument('--kubeconfig',required=True);p.add_argument('--cluster',required=True);p.add_argument('--image',required=True);p.add_argument('--openssl',required=True);p.add_argument('--output',required=True);a=p.parse_args
assert a.cluster.startswith('dr-cks-workload-response-') and '@sha256:' in a.image
assert Path(a.kubeconfig).is_absolute;out=Path(a.output);assert not out.exists
os.umask(0o077);tmp=Path(tempfile.mkdtemp(prefix='dr-cks-mtls-'));ns='dr-mtls-'+uuid.uuid4.hex[:8];records=[];created=False
base=[a.kubectl,'--kubeconfig',a.kubeconfig,'--context','kind-'+a.cluster,'--cache-dir',str(tmp/'cache'),'--request-timeout=15s']
def run(cmd,data=None):
r=subprocess.run(cmd,input=data,text=True,capture_output=True,timeout=60)
if r.returncode:raise RuntimeError('Command failed: '+str(cmd[:3])+': '+r.stderr[:180])
return r.stdout
def k(*args,obj=None):return run(base+list(args),json.dumps(obj) if obj is not None else None)
def create(obj):return k('-n',ns,'create','-f','-',obj=obj)
def get(kind,name):return json.loads(k('-n',ns,'get',kind,name,'-o','json'))
def check(name,value,expected):
assert value==expected,(name,value,expected);records.append(dict(name=name,observed=value,passed=True));print(name,flush=True)
def poll(fn,seconds=150):
deadline=time.monotonic+seconds
while time.monotonic<deadline:
v=fn
if v:return v
time.sleep(1)
raise AssertionError('Deadline exceeded')
def op(*args):return run([a.openssl,*map(str,args)])
def ca(name):
op('req','-x509','-newkey','rsa:2048','-nodes','-days','2','-subj','/CN='+name,'-keyout',tmp/(name+'.key'),'-out',tmp/(name+'.crt'),'-addext','basicConstraints=critical,CA:TRUE','-addext','keyUsage=critical,keyCertSign,cRLSign')
def leaf(name,issuer,san,purpose):
op('req','-new','-newkey','rsa:2048','-nodes','-subj','/CN='+name,'-keyout',tmp/(name+'.key'),'-out',tmp/(name+'.csr'))
ext=tmp/(name+'.ext');ext.write_text('basicConstraints=critical,CA:FALSE\nkeyUsage=critical,digitalSignature,keyEncipherment\nextendedKeyUsage='+purpose+'\nsubjectAltName='+san+'\nsubjectKeyIdentifier=hash\nauthorityKeyIdentifier=keyid,issuer\n')
op('x509','-req','-in',tmp/(name+'.csr'),'-CA',tmp/(issuer+'.crt'),'-CAkey',tmp/(issuer+'.key'),'-CAcreateserial','-days','1','-extfile',ext,'-out',tmp/(name+'.crt'))
def secret(name,files):create(dict(apiVersion='v1',kind='Secret',metadata={'name':name},stringData={key:(tmp/value).read_text for key,value in files.items}))
server=r'''import ssl,socket,json,sys
ctx=ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER);ctx.minimum_version=ssl.TLSVersion.TLSv1_2
ctx.load_cert_chain('/identity/tls.crt','/identity/tls.key');ctx.load_verify_locations('/trust/ca.crt');ctx.verify_mode=ssl.CERT_REQUIRED
s=socket.socket;s.setsockopt(socket.SOL_SOCKET,socket.SO_REUSEADDR,1);s.bind(('0.0.0.0',8443));s.listen(16)
print('ready',flush=True)
while True:
raw,_=s.accept;raw.settimeout(4)
try:
with ctx.wrap_socket(raw,server_side=True) as conn:
cert=conn.getpeercert;allowed=('URI','spiffe://dr.test/ns/funds/sa/reconciler') in cert.get('subjectAltName',)
conn.recv(4096);status=200 if allowed else 403
conn.sendall(('HTTP/1.1 '+str(status)+' Result\r\nContent-Length: 0\r\nConnection: close\r\n\r\n').encode)
except (ssl.SSLError,OSError):raw.close
'''
client=r'''import ssl,socket,json,sys
host,identity,trust,hostname=sys.argv[1:];ctx=ssl.SSLContext(ssl.PROTOCOL_TLS_CLIENT);ctx.minimum_version=ssl.TLSVersion.TLSv1_2
ctx.load_verify_locations('/credentials/'+trust+'.crt')
if identity!='none':ctx.load_cert_chain('/credentials/'+identity+'.crt','/credentials/'+identity+'.key')
try:
with socket.create_connection((host,8443),timeout=4) as raw:
with ctx.wrap_socket(raw,server_hostname=hostname) as s:
s.sendall(b'GET /reconcile HTTP/1.1\r\nHost: funds\r\n\r\n');response=s.recv(4096)
print(json.dumps({'status':int(response.split[1]),'tls':s.version,'peerVerified':True}))
except ssl.SSLCertVerificationError as e:print(json.dumps({'error':'verification','verifyCode':e.verify_code}))
except ssl.SSLError as e:print(json.dumps({'error':'tls','reason':e.reason}))
'''
def pod(name,serving=False):
mounts=[dict(name='code',mountPath='/code',readOnly=True)];vols=[dict(name='code',configMap={'name':'code'})]
if serving:
for key,secretName in [('identity','server'),('trust','client-trust')]:
mounts.append(dict(name=key,mountPath='/'+key,readOnly=True));vols.append(dict(name=key,secret={'secretName':secretName,'defaultMode':0o440}))
else:
mounts.append(dict(name='credentials',mountPath='/credentials',readOnly=True));vols.append(dict(name='credentials',secret={'secretName':'clients','defaultMode':0o440}))
return dict(apiVersion='v1',kind='Pod',metadata={'name':name,'labels':{'app':name}},spec=dict(automountServiceAccountToken=False,terminationGracePeriodSeconds=1,securityContext=dict(runAsNonRoot=True,runAsUser=1000,runAsGroup=1000,fsGroup=1000,seccompProfile={'type':'RuntimeDefault'}),volumes=vols,containers=[dict(name='worker',image=a.image,imagePullPolicy='Never',command=['python','/code/server.py'] if serving else ['python','-c','import time;time.sleep(1800)'],volumeMounts=mounts,securityContext=dict(allowPrivilegeEscalation=False,readOnlyRootFilesystem=True,capabilities={'drop':['ALL']}),resources=dict(requests={'cpu':'10m','memory':'24Mi'},limits={'cpu':'250m','memory':'64Mi'}))]))
def ready(name):poll(lambda:get('pod',name).get('status',{}).get('containerStatuses',[{}])[0].get('ready',False))
def request(target,identity,trust='server-ca',hostname=None):
return json.loads(k('-n',ns,'exec','client','--','python','/code/client.py',get('pod',target)['status']['podIP'],identity,trust,hostname or dns))
def status(target,identity):return request(target,identity).get('status')
def denial(target,identity,trust='server-ca',hostname=None):return request(target,identity,trust,hostname).get('error')
def gen(name):
create(pod(name,True));ready(name);poll(lambda:'ready' in k('-n',ns,'logs',name))
def trust_update(names):
data=''.join((tmp/(n+'.crt')).read_text for n in names)
k('-n',ns,'patch','secret','client-trust','--type=merge','-p',json.dumps({'stringData':{'ca.crt':data}}))
expected=hashlib.sha256(data.encode).hexdigest
code="import hashlib;from pathlib import Path;print(hashlib.sha256(Path('/trust/ca.crt').read_bytes).hexdigest)"
start=time.monotonic;poll(lambda:k('-n',ns,'exec','server-old','--','python','-c',code).strip==expected)
return dict(projected=True,delaySeconds=round(time.monotonic-start,2))
try:
cfg=json.loads(k('config','view','--minify','--raw','-o','json'))['clusters'][0]['cluster'];assert cfg['server'].startswith('https://127.0.0.1:') and not cfg.get('insecure-skip-tls-verify');cfg=None
versions=json.loads(k('version','-o','json'));assert versions['serverVersion']['gitVersion']=='v1.37.0' and versions['clientVersion']['gitVersion']=='v1.37.1'
k('create','namespace',ns);created=True;k('label','namespace',ns,'pod-security.kubernetes.io/enforce=restricted','pod-security.kubernetes.io/enforce-version=v1.35')
dns='funds.'+ns+'.svc'[ca(n) for n in ['server-ca','old-ca','new-ca','rogue-ca']]
leaf('server','server-ca','DNS:'+dns,'serverAuth')
for name,issuer,identity in [('old','old-ca','reconciler'),('new','new-ca','reconciler'),('unauthorized','old-ca','reporter'),('rogue','rogue-ca','reconciler')]:leaf(name,issuer,'URI:spiffe://dr.test/ns/funds/sa/'+identity,'clientAuth')
secret('server',{'tls.crt':'server.crt','tls.key':'server.key'});secret('client-trust',{'ca.crt':'old-ca.crt'})
files={n+suffix:n+suffix for n in ['old','new','unauthorized','rogue'] for suffix in ['.crt','.key']};files.update({'server-ca.crt':'server-ca.crt','rogue-ca.crt':'rogue-ca.crt'});secret('clients',files)
create(dict(apiVersion='v1',kind='ConfigMap',metadata={'name':'code'},data={'server.py':server,'client.py':client}));create(pod('client'));ready('client');gen('server-old')
check('restricted-pods-running',all(get('pod',n)['spec']['securityContext']['runAsUser']==1000 for n in ['client','server-old']),True)
runtime=json.loads(k('-n',ns,'exec','client','--','python','-c',"import sys,ssl,platform,json;print(json.dumps({'python':platform.python_version,'openssl':ssl.OPENSSL_VERSION,'kernel':platform.release,'architecture':platform.machine}))"));images={n:get('pod',n)['status']['containerStatuses'][0]['imageID'] for n in ['client','server-old']}
check('old-authorized-client-accepted',status('server-old','old'),200)
check('missing-client-certificate-rejected',denial('server-old','none'),'tls')
check('untrusted-client-certificate-rejected',denial('server-old','rogue'),'tls')
check('trusted-unauthorized-client-denied',status('server-old','unauthorized'),403)
check('wrong-server-hostname-rejected',denial('server-old','old',hostname='different.dr.test'),'verification')
check('wrong-server-trust-rejected',denial('server-old','old',trust='rogue-ca'),'verification')
check('new-client-before-overlap-rejected',denial('server-old','new'),'tls')
overlap=trust_update(['old-ca','new-ca']);check('overlap-trust-projected',overlap['projected'],True)
check('loaded-context-still-rejects-new-client',denial('server-old','new'),'tls')
gen('server-overlap');check('replacement-context-accepts-old-client',status('server-overlap','old'),200);check('replacement-context-accepts-new-client',status('server-overlap','new'),200)
retirement=trust_update(['new-ca']);check('new-only-trust-projected',retirement['projected'],True)
check('old-generation-still-accepts-old-client',status('server-old','old'),200)
gen('server-new');check('new-generation-rejects-old-client',denial('server-new','old'),'tls');check('new-generation-accepts-new-client',status('server-new','new'),200)
check('overlap-generation-still-accepts-old-client',status('server-overlap','old'),200)
for n in ['server-old','server-overlap']:k('-n',ns,'delete','pod',n,'--wait=true','--timeout=40s')
remaining=json.loads(k('-n',ns,'get','pods','-o','json'))['items'];check('old-generations-removed',sorted(x['metadata']['name'] for x in remaining),['client','server-new'])
check('accepted-flow-survives-retirement',status('server-new','new'),200)
check('retired-client-still-denied',denial('server-new','old'),'tls')
assert len(records)==20
finally:
if created:k('delete','namespace',ns,'--wait=true','--timeout=45s')
shutil.rmtree(tmp)
report=dict(executedAt=datetime.datetime.now(datetime.timezone.utc).isoformat,scriptSha256=hashlib.sha256(Path(__file__).read_bytes).hexdigest,versions=versions,runtime=runtime,issuerOpenSSL=op('version').strip,images=images,examVersion='1.35',policyVersion='v1.35',observations=records,projectionDelays={'overlap':overlap,'retirement':retirement},passed=True,cleanup=dict(namespaceRemoved=True,privateMaterialRemoved=not tmp.exists,realCredentialsUsed=False),scope='Actual application mTLS between Pods using synthetic Secrets, hostname and trust failures, authorization403, trust overlap and replacement server generations. No mesh,CNI policy,sandbox,encryption-at-rest,revocation service,active-session termination,production recovery or full practical mock executed.')
out.write_text(json.dumps(report,indent=2)+'\n');print('PASS',len(records))
The mounted bundle includes the new CA, but the old server still rejects it; a new generation accepts both clients during overlap.
Common pitfalls
Confusing projection with adoption, authentication with authorization, one replica with the whole service or application mTLS with a tested mesh.
Related topics: Pod Security and recovery · Secrets and key rotation · mTLS and authorization · Isolation and handover
Recovery is established when intended consumers work and access boundaries remain observable across relevant generations.
Reference: CKS domains and exam details · Kubernetes v1.35; current six-domain CKS outline