Define what is protected
A reconciliation application uses a password held in a Secret. Three distinct questions arise: who can obtain the value through the API, how persistent storage represents that value and whether the database still accepts the password. Associate each question with its control and observation. An authorized administrator can still read content after at-rest encryption is enabled. Resource use depends on that access, and the read alone does not establish a storage failure. In the laboratory, observe a write with identity and then one with aescbc. Compare the prefix and presence of the synthetic marker in etcd with retrieval of the same value through the API. The base64 representation in data may remain identical across both reads. Explain this difference to the project owner: a kubectl get capture does not replace evidence from persistent storage, and neither observation establishes that an exposed credential has stopped working in an external system. The exercise uses aescbc to make migration phases visible. Kubernetes 1.35 documentation does not recommend it for new production choices because of padding-oracle risk. Do not copy this selection as a production architecture. Assess a suitable provider, key management and access, recovery dependencies and the threat model. The operational comparison observed here is neither a cryptographic evaluation nor design approval. Keep the learning objective explicit when handing this exercise to another engineer.
Prepare a hypothesis and ordered configuration
Before execution, write a prediction for every state. With identity first, expect new writes in plaintext. With aescbc first and key1 as its first key, expect new writes under key1. An earlier object may retain its previous representation until rewritten. If a prediction fails, confirm adopted configuration and the observed object before modifying further components. This discipline reduces contradictory changes during a short maintenance window. Resource selection also bounds the conclusion. Protecting secrets does not automatically cover a ConfigMap in the same namespace. When using wildcards, review entry ordering and specific exclusions, including invalid combinations. The EncryptionConfiguration reference describes that structure. Draw a small matrix containing resource, applicable entry, first provider and reader for existing data. Use it to detect an exception placed after a broader entry or an incorrect assumption of protection by namespace. In the fictional funds case, the team discovers a sensitive connection string in a ConfigMap. The decision includes reviewing where the value belongs and who can read it, alongside encryption scope. Do not classify the control as fully effective because one Secret was observed. Also record what was not exercised: the laboratory executes Secret selection; wildcards and exclusions are documentation-analysis exercises in this lesson. A useful review identifies both the intended boundary and a concrete observation that could disprove the team's assumption about it.
Run the laboratory and record its boundaries
The attached code creates a temporary kind cluster with one control plane, a localhost API and a private kubeconfig. It needs Python 3, Docker, kind 0.33.0, kubectl 1.35.8 and the Kubernetes 1.35.8 image pinned by the digest in the code. Arguments --kind, --kubectl and --report select executables and report destination. It does not use the operator's normal context. Random keys and four Secret values belong exclusively to this exercise; reports retain observations and identifiers without publishing key material. The exercise observes six automatic reloads and 32 checks: initial writing, enablement, unmigrated older data, a second key not yet writing, promotion, premature-retirement failure, recovery and retirement after migration. Watch cache is disabled to observe storage-read failure without a previously decrypted cached response. This choice bounds the test: it does not establish normal cache behavior, performance or availability in a redundant control plane. Before execution, predict which Secrets remain readable at each phase. Then compare predictions with the report and explain differences. Final cleanup removes only the cluster and temporary material created by the script. Both recorded final runs passed the same 32 checks. Neither used production data. This exercise does not execute KMS, snapshot restoration, distributed rotation or a full CKS practical mock. Use the evidence to support the exact behavior observed, while keeping those other capabilities open for separate assessment.
Migrate without losing concurrent changes
The first Secret retains its readable marker in etcd after encryption becomes the first provider. Creating another Secret already produces a key1 prefix. This comparison demonstrates why a new-write check does not close work on earlier data. Rewriting the first object changes representation while preserving the API-returned value. Record these as separate observations: storage protection and content integrity. With concurrent writers, a read followed by PUT may encounter a different resource version. A 409 Conflict requires fetching current state and reassessing the write while preserving legitimate changes. Do not turn the procedure into restoration of an old copy of every credential. The API reference explains resourceVersion and lost-update detection. The laboratory does not deliberately introduce concurrency; conflict handling is a separate decision exercise. Define scope before removing compatibility. If EncryptionConfiguration protects Secrets in every namespace, migrating funds alone leaves a gap. The script rewrites all Secrets in its disposable cluster, including those generated for that cluster itself, and checks their prefix before the final phase. Do not generalize that authorization to shared clusters. In operations, identify owners, impact, change control and evidence required for the actual dataset. A completed application checklist must not be reused as proof that unrelated namespaces no longer depend on an older reader.
Separate key introduction, promotion and retirement
During introduction, key2 follows key1. A new Secret named staged still uses key1. During promotion, key2 moves first and Secret new2 uses it. Meanwhile, an object previously written under key1 still depends on the earlier reader. These three observations explain the change without claiming every object changed simultaneously. With automatic reload, file modification time does not confirm adoption. The script waits for apiserver_encryption_config_controller_automatic_reload_last_timestamp_seconds to advance and then makes read or write observations. Waiting can approach a minute per phase. Set a limit and record failure rather than continuing as though reload occurred. Six phases deliberately make this exercise slower than syntax validation. In a control plane with multiple instances, establish reading capability on every instance before promoting new writes. A query through a load balancer does not necessarily identify the instance still using older configuration. Build an instance-level plan and progression criterion. The lesson applies documentation to this problem, but the local cluster has only one API server. In handover, describe the topology actually observed so another operator does not interpret local evidence as a high-availability exercise. Record which actions remain conceptual and which were repeated against the pinned runtime, allowing the next assessment to target the missing behavior.
Recover from premature retirement
The laboratory removes key1 while ciphertext still depends on it. Reading the older Secret fails; a key2 Secret remains readable. The hash of the earlier bytes stays unchanged. The incident demonstrates read unavailability rather than automatic disappearance of the object. That distinction changes response: preserving recoverable data and restoring its reader may be better than deleting and recreating incomplete resources. The next phase restores key1 as a reader after key2. The earlier read succeeds again. Only then does the script rewrite the Secret dataset and verify key2 use before removing key1 and identity. Relate failure, ciphertext preservation and recovery in the report. Do not conclude that every Secret failure has this cause; the case deliberately establishes known state to isolate the behavior. During a fictional fund closing, unavailability may affect jobs with different dependencies. APS should identify affected jobs, retain processing state and confirm resumption criteria with business and security. One readable Secret does not establish that every job can be replayed without duplication. A temporary alternative may be holding the affected batch with communicated impact. This operational sequence is educational and does not describe internal BNP Paribas procedures. Ask the next engineer to explain both the technical recovery condition and the separate business condition required before processing resumes.
Relate recovery, KMS and credential exposure
An older snapshot may contain data still protected by a key retired from current state. File integrity does not establish the ability to read that historical state. As a design exercise, relate every recovery point to configuration and authorized access to required material. Merely renaming key2 to key1 does not recreate the earlier key. The laboratory does not restore snapshots: this dependency is analyzed conceptually and needs its own recovery exercise. KMS also requires distinguishing health from coverage. Kubernetes documentation recommends KMS v2 where possible and describes integration and health checks. A healthy provider does not inventory or rewrite every older object. In a pilot, request evidence of adoption, new writes, migration and recovery, with each observation's limits. Dependency unavailability may have different effects according to caches and data state, which this aescbc cluster did not exercise. Finally, rotating a storage key while retaining a password value does not revoke that password. In the fictional case, coordinate change in the system accepting it, adoption by consumers and verification of the previous value. Review existing sessions and continuity with their owners. Do not close an exposed-credential incident merely because the Secret prefix changed from key1 to key2. Storage protection and credential lifecycle have connected operational consequences, while each needs evidence from its own system.
Hand over evidence the next team can use
Produce a matrix containing initial state, change, observation, interpretation and next decision. Include at least five rows: older plaintext after enablement, key2 introduced without writing, key2 promoted with remaining key1 data, read failure after retirement and recovery after restoration. In every row, identify whether evidence comes from the API, storage or adopted configuration. Avoid treating absence of an error as proof of all three. To close the exercise, another person should be able to explain why the first retirement failed and the second succeeded without receiving secret values. Provide versions, image digest, code hash, observed phases, results and cleanup. Add pending capabilities: redundancy, deliberate conflicts, normal caching, KMS and snapshot restoration. Two automatic runs improve repeatability of observation; they do not constitute independent specialist review or certification. Summary: define scope, distinguish writer from readers, confirm adoption, migrate while preserving values and remove compatibility only after observing the relevant dataset. When an application uses an exposed credential, also address that credential in its external system. Connect these steps to RBAC, auditing, disaster recovery, change management and incident response. Use the module questions to justify the next action and explain why each alternative fails in its stated context. Retain unresolved dependencies in the handover so completion of this exercise does not silently close work that still requires evidence.
#!/usr/bin/env python3
"""Original disposable Kubernetes 1.35 encryption-at-rest exercise.
Creates and deletes its own kind cluster. Only generated keys and synthetic Secrets.
Requires Docker, kind 0.33 and kubectl 1.35.8. Never uses the default kubeconfig.
Usage: python3 run.py --kind /path/kind --kubectl /path/kubectl --report /path/new.json
"""
import argparse,base64,datetime,hashlib,json,os,pathlib,re,shutil,subprocess,tempfile,time,uuid
IMAGE='kindest/node:v1.35.8@sha256:07b2536e30b803ed61d1677a79df6115f798ce64c80f9e22f6ed45afd09323c0'
PREFIX=b'k8s:enc:aescbc:v1:'
def main:
ap=argparse.ArgumentParser;ap.add_argument('--kind',required=True);ap.add_argument('--kubectl',required=True);ap.add_argument('--report',type=pathlib.Path,required=True);a=ap.parse_args;assert not a.report.exists
os.umask(0o077);work=pathlib.Path(tempfile.mkdtemp(prefix='dr-cks-encryption-'));enc=work/'enc'enc.mkdir;cfg=enc/'config.json'kubeconfig=work/'kubeconfig'name='dr-cks-encryption-'+uuid.uuid4.hex[:10];node=name+'-control-plane'ns='dr-encryption'records=[];phases=[];started=False
keys={x:base64.b64encode(os.urandom(32)).decode for x in ['key1','key2']};marker='DR_ORIGINAL_SYNTHETIC_VALUE_20261007'
report=dict(startedAt=datetime.datetime.now(datetime.timezone.utc).isoformat,scriptSha256=hashlib.sha256(pathlib.Path(__file__).read_bytes).hexdigest,nodeImage=IMAGE,cluster=name,observations=records,phases=phases,syntheticOnly=True,actualCluster=True,actualEtcdInspection=True,actualEncryptionRotation=True,HA=False,KMS=False,backupRestore=False,watchCache=False,fullPracticalMock=False,independentVerification=False)
def run(cmd,data=None,timeout=25,check=True):
r=subprocess.run(cmd,input=data,capture_output=True,text=True,timeout=timeout)
if check and r.returncode:raise RuntimeError('Command failed ('+str(r.returncode)+'): '+str(cmd[:2])+' '+r.stderr[:160])
return r
def k(*args,obj=None,check=True):return run([a.kubectl,'--kubeconfig',str(kubeconfig),'--context','kind-'+name,'--cache-dir',str(work/'cache'),'--request-timeout=15s',*args],json.dumps(obj) if obj is not None else None,check=check)
def observe(label,value,expected):
records.append(dict(name=label,observed=value,expected=expected,passed=value==expected));print(label,flush=True);assert value==expected,(label,value,expected)
def config(order,identity=False,identity_first=False):
providers=[{'aescbc':{'keys':[{'name':n,'secret':keys[n]} for n in order]}}]
if identity:providers=([{'identity':{}}]+providers) if identity_first else (providers+[{'identity':{}}])
d=dict(apiVersion='apiserver.config.k8s.io/v1',kind='EncryptionConfiguration',resources=[{'resources':['secrets'],'providers':providers}]);tmp=enc/'pending.json'tmp.write_text(json.dumps(d));tmp.chmod(0o600);tmp.replace(cfg)
def metric:
text=k('get','--raw','/metrics').stdout;values=[float(line.rsplit(' ',1)[1]) for line in text.splitlines if line.startswith('apiserver_encryption_config_controller_automatic_reload_last_timestamp_seconds')];return max(values,default=0)
def wait(fn,timeout=150):
end=time.monotonic+timeout
while time.monotonic<end:
try:
x=fn
if x:return x
except (RuntimeError,subprocess.TimeoutExpired):pass
time.sleep(2)
raise AssertionError('Timed out waiting for explicit observed state')
def switch(label,order,identity=False):
prior=metric;config(order,identity);began=time.monotonic;observed=wait(lambda:metric>prior);phases.append(dict(name=label,keys=order,identityFallback=identity,reloadMetricAdvanced=observed,elapsedSeconds=round(time.monotonic-began,2)));observe(label+'-reload-observed',observed,True)
def secret(n):return dict(apiVersion='v1',kind='Secret',metadata={'name':n,'namespace':ns},stringData={'sample':marker+'_'+n})
def create(n):k('create','-f','-',obj=secret(n))
def get(n,check=True):return k('-n',ns,'get','secret',n,'-o','json',check=check)
def value(n):return base64.b64decode(json.loads(get(n).stdout)['data']['sample']).decode
def raw(n,namespace=ns):
r=k('-n','kube-system','exec','etcd-'+node,'--','etcdctl','--endpoints=https://127.0.0.1:2379','--cacert=/etc/kubernetes/pki/etcd/ca.crt','--cert=/etc/kubernetes/pki/etcd/healthcheck-client.crt','--key=/etc/kubernetes/pki/etcd/healthcheck-client.key','get','/registry/secrets/'+namespace+'/'+n,'-w','json')
data=json.loads(r.stdout);assert data.get('count')==1;return base64.b64decode(data['kvs'][0]['value'])
def stored(n,key):return raw(n).startswith(PREFIX+key.encode+b':')
def rewrite(n):k('replace','-f','-',obj=json.loads(get(n).stdout))
try:
config(['key1'],True,True)
patch=dict(apiVersion='kubeadm.k8s.io/v1beta3',kind='ClusterConfiguration',apiServer=dict(extraArgs={'encryption-provider-config':'/etc/kubernetes/enc/config.json','encryption-provider-config-automatic-reload':'true','watch-cache':'false'},extraVolumes=[dict(name='dr-encryption',hostPath='/etc/kubernetes/enc',mountPath='/etc/kubernetes/enc',readOnly=True,pathType='Directory')]))
kindcfg=dict(kind='Cluster',apiVersion='kind.x-k8s.io/v1alpha4',networking={'apiServerAddress':'127.0.0.1'},nodes=[dict(role='control-plane',extraMounts=[dict(hostPath=str(enc),containerPath='/etc/kubernetes/enc',readOnly=True)],kubeadmConfigPatches=[json.dumps(patch)])]);(work/'kind.json').write_text(json.dumps(kindcfg))
started=True;print('creating-owned-cluster',name,flush=True)
run([a.kind,'create','cluster','--name',name,'--image',IMAGE,'--config',str(work/'kind.json'),'--kubeconfig',str(kubeconfig),'--wait','120s'],timeout=240)
report['versions']=json.loads(k('version','-o','json').stdout);observe('server-version',report['versions']['serverVersion']['gitVersion'],'v1.35.8');observe('client-version',report['versions']['clientVersion']['gitVersion'],'v1.35.8')
observe('private-config-mode',oct(cfg.stat.st_mode & 0o777),'0o600');k('create','namespace',ns)
create('legacy');observe('identity-stores-plaintext',(marker+'_legacy').encode in raw('legacy'),True);observe('identity-api-readable',value('legacy'),marker+'_legacy')
switch('enable-key1',['key1'],True);create('new1');observe('new-secret-key1',stored('new1','key1'),True);observe('new-secret-marker-hidden',marker.encode in raw('new1'),False);observe('new-secret-api-readable',value('new1'),marker+'_new1');observe('legacy-not-auto-migrated',(marker+'_legacy').encode in raw('legacy'),True)
rewrite('legacy');observe('legacy-rewritten-key1',stored('legacy','key1'),True);observe('legacy-preserved-after-migration',value('legacy'),marker+'_legacy')
switch('stage-key2',['key1','key2'],True);create('staged');observe('second-key-not-writer',stored('staged','key1'),True)
switch('promote-key2',['key2','key1'],True);create('new2');observe('new-secret-key2',stored('new2','key2'),True);observe('old-ciphertext-remains-key1',stored('legacy','key1'),True);observe('old-key-still-decrypts',value('legacy'),marker+'_legacy')
old_cipher=hashlib.sha256(raw('legacy')).hexdigest;switch('premature-retirement',['key2'],True)
broken=get('legacy',False);observe('missing-key-read-fails',broken.returncode!=0,True);observe('missing-key-storage-error',any(x in broken.stderr for x in ['InternalError','unable to transform','no matching key']),True);observe('failed-read-preserves-ciphertext',hashlib.sha256(raw('legacy')).hexdigest,old_cipher);observe('current-key-readable-during-failure',value('new2'),marker+'_new2')
switch('restore-old-reader',['key2','key1'],True);observe('restored-key-recovers-read',value('legacy'),marker+'_legacy')
inventory=json.loads(k('get','secrets','-A','-o','json').stdout)['items']
for obj in inventory:k('replace','-f','-',obj=obj)
report['rewrittenSecretCount']=len(inventory)
observe('all-disposable-cluster-secrets-rewritten',all(raw(obj['metadata']['name'],obj['metadata']['namespace']).startswith(PREFIX+b'key2:') for obj in inventory),True)
observe('all-exercise-secrets-rewritten',all(stored(n,'key2') for n in ['legacy','new1','staged','new2']),True)
switch('retire-after-rewrite',['key2'],False)
observe('cluster-secret-list-readable',len(json.loads(k('get','secrets','-A','-o','json').stdout)['items'])==len(inventory),True)
observe('all-exercise-secrets-readable',all(value(n)==marker+'_'+n for n in ['legacy','new1','staged','new2']),True)
observe('all-exercise-secrets-key2',all(stored(n,'key2') for n in ['legacy','new1','staged','new2']),True)
observe('no-synthetic-marker-in-stored-ciphertext',all(marker.encode not in raw(n) for n in ['legacy','new1','staged','new2']),True)
report['scopeNote']='One disposable API server, watch cache explicitly disabled to observe storage read failures, local aescbc keys and four original synthetic Secrets. Final rewrite covers every Secret in this disposable cluster, including its generated bootstrap material. Six observed automatic reloads. No HA rollout,KMS,backup restoration,production secrets or real incident recovery.'report['passed']=True
except BaseException as exc:report['passed']=False;report['error']=str(exc);raise
finally:
if started:run([a.kind,'delete','cluster','--name',name,'--kubeconfig',str(kubeconfig)],timeout=90,check=False)
remaining=run(['docker','ps','-a','--filter','label=io.x-k8s.kind.cluster='+name,'--format','{{.Names}}'],check=False).stdout.strip;shutil.rmtree(work)
report['cleanup']=dict(ownedClusterRemoved=not remaining,temporaryKeysRemoved=not work.exists,dedicatedKubeconfigRemoved=not kubeconfig.exists,realCredentialsUsed=False,defaultKubeconfigModified=False);report['finishedAt']=datetime.datetime.now(datetime.timezone.utc).isoformat;a.report.parent.mkdir(parents=True,exist_ok=True);a.report.write_text(json.dumps(report,indent=2)+'\n')
print(json.dumps({'passed':report['passed'],'observations':len(records),'report':str(a.report)}))
if __name__=='__main__':main
Removing key1 before rewriting older objects prevents reads; restoring the reader recovers preserved values.
Common pitfalls
Treating base64 as confidentiality evidence, retiring readers prematurely or confusing storage rotation with password revocation.
Related topics: RBAC and Secrets · Migration and concurrency · Disaster recovery · Change management
One protected new write does not establish migration or recovery of all earlier data.
Reference: CKS domains and exam details · Kubernetes v1.35; current six-domain CKS outline