Reduce exposure while preserving operation
In a fictional funds-platform modernization project, the node pool contains tools installed during earlier incidents. Some have no owner, and a legacy exporter is still listening. Before removing components, identify purpose, dependencies, running processes, future activation and the recovery path. No visible use during one hour does not prove a component is unnecessary at month-end close. Turn the decision into a change with an owner, a window, a business test and a fallback criterion. The objective is to reduce known exposure while retaining approved functions. Distinguish installation, activation and execution. Systemctl disable does not itself terminate a running service. Stopping it also does not establish removal of an activation source, such as a socket unit still requiring investigation. Review actual units and endpoints. In an image-managed pool, compare the baseline with real node state and close diagnostic exceptions. Red Hat image-mode documentation illustrates this drift problem; it does not mean every project host uses that product. Investigation should end with an updated inventory and a support path that does not depend on forgotten local changes.
Read effective process identity
A hardening configuration must correspond to the process running the application. Record UID, GID, and effective, permitted and bounding capability sets instead of assuming everything is restricted because the username differs from root. A nonzero UID is not universal proof that capabilities are absent. In this lesson’s exercise, the process has UID and GID 1000 and all three observed sets are empty. An attempt to switch to UID zero is rejected. That outcome belongs to this combination of identity and controls; do not attribute it in isolation to one flag. No_new_privs addresses a specific problem: preventing certain privilege gains when executing a new program. It does not erase all existing capabilities. A privilege-reduction review must identify rights already present and those obtainable through permitted paths. The lab queries the flag and checks its inheritance in a new process. Connect this test to a workplace need: a reconciliation agent that only reads approved files should not receive administrative rights to resolve a permission error before its actual access requirement has been checked.
Choose controls for the right resource
Seccomp filters calls into the kernel; it does not replace every access-control mechanism. A classic filter observes syscall numbers and arguments but does not dereference a pointer to compare arbitrary pathname text. If the requirement is to restrict writes to specific paths, also assess permissions, mounts and suitable resource-access policies. AppArmor and SELinux have their own models requiring distribution and validation. Do not present a seccomp rule permitting openat as a guarantee that only an approved file can be opened. A read-only root also has a specific scope. A separate volume can remain writable if its mount and permissions allow it. For a report-generating job, define temporary storage, result publication and which image areas must stay immutable. In the exercise, a write attempt under /var/tmp encounters EROFS, while /tmp is a writable temporary mount. This comparison helps diagnose failures without disabling protection for the whole root filesystem. Acceptance should test required functionality and an operation the control intends to reject, retaining the path and identity used in each observation.
Distribute profiles and restrict management access
A Localhost profile referenced by a Pod must be available on the node where its container starts. If pool A has AppArmor profile close-report and pool B does not, a successful test in A does not cover B. Define profile distribution and placement criteria, and verify the profile actually applied. Kubernetes 1.35 documentation explains that the scheduler does not automatically discover loaded AppArmor profiles. A pool or node-image change must retest this dependency. This lesson’s local lab does not execute AppArmor or replace that cluster test. Review exceptions that can override other fields. A privileged container does not necessarily retain the seccomp confinement its manifest appears to request. Functional success in that mode does not establish that the earlier profile was wrong or that the configuration is secure. Separately, distinguish application exposure from management-plane exposure. A public application does not need a public kubelet. Define authorized sources, administrative paths and denial evidence from an unauthorized source. TLS, authentication and network controls are separate layers; an outcome at one layer does not automatically establish enforcement of the others.
Prepare the exercise and predict outcomes
Save the code below as run.py in its own directory and choose a new report path. The exercise requires host Python, Docker with a Linux aarch64 runtime, and the image identified by the code’s digest already available locally. The script uses --pull=never and refuses to overwrite a report. Recorded execution used Docker 29.1.3, kernel 6.12.54-linuxkit and Python 3.13.16 inside the container. These values describe the observed trial; they do not establish reproduction of the official CKS environment, whose page states Kubernetes 1.35. Run python3 run.py --output evidence.json. The randomly named container has no network, uses UID/GID 1000, dropped capabilities, a read-only root and a bounded temporary mount. It retains the runtime’s built-in seccomp filter. Only the exercise file is mounted from the host, and that mount is read-only. CPU, memory and process counts are bounded. Before execution, predict root-write, tmp-write, no_new_privs query and UID-change outcomes. Also write what each result can establish. A network setting declared in arguments is not a firewall-testing campaign, which this exercise does not perform.
Follow an additional restriction
The program first observes its own process state and a working query. It then installs an original BPF filter rejecting only prctl with the PR_GET_NO_NEW_PRIVS operation while retaining runtime filters. The filter checks the aarch64 audit architecture first, then the syscall number and operation argument. Offsets and constants were checked against Linux 6.12 UAPI headers. The implementation is limited to its declared architecture; the same numbers are not a universal interface across ABIs. An ABI mismatch should stop the experiment instead of producing a conclusion about another call. The objective is to observe a controlled difference. The query returns 1 before the filter and returns -1 with EPERM afterward. Meanwhile, /proc/self/status still reports NoNewPrivs: 1. The flag was not unset: the query operation was denied. The filter count increases from one to two in the recorded execution. An unrelated read continues to work. Explain these four observations together. Installing an additional filter that permits other calls does not remove higher-precedence denials already imposed by the runtime.
Inheritance,evidence and cleanup
The exercise starts a new Python process before and after adding the filter. In the second child execution, the query remains denied and the flag remains set. This observation covers the creation-and-execution path used by the script. It does not establish synchronized filter updates across an application’s existing threads or validate every possible call. Installation through seccomp without the TSYNC flag affects the calling thread; accepting another concurrency model requires its own tests. This distinction prevents transferring a local conclusion to a production agent with a different architecture. The report retains 14 observations, versions and the executed file’s hash. The script removes its own container and confirms absence. The temporary mount disappears with that container; host profiles are unchanged. Two final executions used identical code and both completed. The test does not execute Kubernetes, AppArmor, SELinux, systemd or host firewall rules. It also does not establish complete isolation, exploit resistance or load performance. When handing over evidence, associate each claim with its supporting outcome and explicitly retain areas that need their own lab.
Decide recovery and prepare RUN
Consider a closing batch that starts failing after a release changes UID, capabilities and seccomp. There are 25 minutes until the contingency decision and an approved previous version exists. Collect the rejected operation, effective identity and configuration differences. Use agreed change criteria to separate business recovery from investigation that needs more time. Running privileged can make the job finish while undoing the control the team intended to introduce. That outcome neither identifies the cause by itself nor replaces security acceptance. RUN handover should contain expected behavior, errors requiring diagnosis and the escalation point. For each mechanism, define a functional test and an expected denial: writes to the approved volume and rejection on root; a necessary call and a restricted call; authorized management access and a rejected source. Record who maintains profiles, images and rules, and which changes require renewed validation. A useful summary is a set of verifiable decisions with scope and owners. This lesson’s banking cases are original and fictional; they do not describe internal BNP Paribas procedures.
"""Original DR Linux aarch64 hardening lab. Run on host with --output; container uses --inside."""
import sys
def probe:
"""Original Linux aarch64 experiment: add a restrictive filter to the current thread."""
import ctypes,errno,json,os,platform,subprocess,sys
from pathlib import Path
assert platform.system=='Linux' and platform.machine=='aarch64'
libc=ctypes.CDLL(None,use_errno=True);libc.prctl.restype=ctypes.c_int
PR_GET_NO_NEW_PRIVS=39;SECCOMP_SET_MODE_FILTER=1;SYS_SECCOMP=277;libc.syscall.restype=ctypes.c_long
def prctl(option,arg=0):
ctypes.set_errno(0);result=libc.prctl(ctypes.c_int(option),ctypes.c_ulong(arg),ctypes.c_ulong(0),ctypes.c_ulong(0),ctypes.c_ulong(0));return dict(result=result,errno=ctypes.get_errno)
def status:
raw=dict(line.split(':',1)for line in Path('/proc/self/status').read_text.splitlinesif ':'in line)
return {k:raw[k].stripfor k in ['Uid','Gid','CapEff','CapPrm','CapBnd','NoNewPrivs','Seccomp','Seccomp_filters']}
if '--child'in sys.argv:
print(json.dumps(dict(query=prctl(PR_GET_NO_NEW_PRIVS),status=status)));sys.exit(0)
records=[]
def check(name,condition,observed):
assert condition,(name,observed);records.append(dict(name=name,passed=True,observed=observed))
initial=status;query=prctl(PR_GET_NO_NEW_PRIVS)
check('non-root-identity',os.getuid==1000 and os.getgid==1000,dict(uid=os.getuid,gid=os.getgid))
check('capability-sets-empty',all(int(initial[k],16)==0 for k in ['CapEff','CapPrm','CapBnd']),{k:initial[k]for k in ['CapEff','CapPrm','CapBnd']})
check('no-new-privileges-set',initial['NoNewPrivs']=='1'and query==dict(result=1,errno=0),query)
check('runtime-filter-present',initial['Seccomp']=='2'and int(initial['Seccomp_filters'])>=1,initial)
# /var/tmp has writable DAC permissions in this image; EROFS identifies the mount constraint.
try:Path('/var/tmp/dr-root-write').write_text('synthetic');root_errno=0
except OSError as err:root_errno=err.errno
check('root-filesystem-read-only',root_errno==errno.EROFS,dict(errno=root_errno))
path=Path('/tmp/dr-writable');path.write_text('synthetic');check('temporary-volume-writable',path.read_text=='synthetic',True);path.unlink
try:os.setuid(0);uid_errno=0
except OSError as err:uid_errno=err.errno
check('uid-change-rejected',uid_errno==errno.EPERM and os.getuid==1000,dict(errno=uid_errno,uid=os.getuid))
child=json.loads(subprocess.check_output([sys.executable,'-B',__file__,'--inside','--child'],text=True,timeout=10));check('baseline-child-inherits-nnp',child['query']==dict(result=1,errno=0)and child['status']['NoNewPrivs']=='1',child)
class Filter(ctypes.Structure):_fields_=[('code',ctypes.c_ushort),('jt',ctypes.c_ubyte),('jf',ctypes.c_ubyte),('k',ctypes.c_uint)]
class Program(ctypes.Structure):_fields_=[('len',ctypes.c_ushort),('filter',ctypes.POINTER(Filter))]
# Original BPF: reject unexpected ABI, then deny only prctl(PR_GET_NO_NEW_PRIVS).
# Linux6.12 aarch64: audit arch0xc00000b7, prctl syscall167; nr@0,arch@4,args[0]@16.
rows=[(0x20,0,0,4),(0x15,1,0,0xc00000b7),(0x06,0,0,0x80000000),
(0x20,0,0,0),(0x15,0,3,167),(0x20,0,0,16),(0x15,0,1,39),
(0x06,0,0,0x00050000|errno.EPERM),(0x06,0,0,0x7fff0000)]
filters=(Filter*len(rows))(*(Filter(*row)for row in rows));program=Program(len(rows),filters)
ctypes.set_errno(0);result=libc.syscall(ctypes.c_long(SYS_SECCOMP),ctypes.c_uint(SECCOMP_SET_MODE_FILTER),ctypes.c_uint(0),ctypes.byref(program))
check('additional-filter-installed',result==0,dict(result=result,errno=ctypes.get_errno))
after=status;blocked=prctl(PR_GET_NO_NEW_PRIVS)
check('query-now-denied',blocked==dict(result=-1,errno=errno.EPERM),blocked)
check('flag-remains-set',after['NoNewPrivs']=='1',after['NoNewPrivs'])
check('filter-count-increased',int(after['Seccomp_filters'])==int(initial['Seccomp_filters'])+1,dict(before=initial['Seccomp_filters'],after=after['Seccomp_filters']))
child=json.loads(subprocess.check_output([sys.executable,'-B',__file__,'--inside','--child'],text=True,timeout=10));check('exec-child-inherits-filter',child['query']==blocked and child['status']['NoNewPrivs']=='1',child)
check('unrelated-read-still-works',Path('/etc/os-release').read_text.startswith('NAME='),True)
assert len(records)==14
print(json.dumps(dict(python=platform.python_version,kernel=platform.release,architecture=platform.machine,observations=records)))
def run:
"""Run only a newly named, resource-limited Docker container; keep runtime seccomp."""
import argparse,datetime,hashlib,json,subprocess,uuid
from pathlib import Path
p=argparse.ArgumentParser;p.add_argument('--output',required=True);a=p.parse_args;out=Path(a.output);assert not out.exists;here=Path(__file__).resolve.parent
image='python@sha256:2d9aefe2fef018a7eb2c13064c89c71929800fd2e5dccdbf52ea5da5bb8d929a'name='dr-cks-system-'+uuid.uuid4.hex[:12]
def docker(*args,ok=True,timeout=30):
r=subprocess.run(['docker',*args],capture_output=True,text=True,timeout=timeout)
if ok and r.returncode:raise RuntimeError(r.stderr[:600])
return r
info=json.loads(docker('info','--format','{{json.}}').stdout);assert info['Architecture']=='aarch64'assert any('seccomp'in x for x in info['SecurityOptions'])
try:
args=['run','--rm','--pull=never','--name',name,'--label','dr.lesson=cks-system','--network=none','--read-only','--user','1000:1000','--cap-drop=ALL','--security-opt=no-new-privileges=true','--security-opt=seccomp=builtin','--pids-limit=32','--memory=64m','--cpus=0.5','--tmpfs','/tmp:rw,noexec,nosuid,nodev,size=1m,mode=1777','--mount','type=bind,src='+str(Path(__file__).resolve)+',dst=/lesson/run.py,readonly','--entrypoint','python',image,'-B','/lesson/run.py','--inside']
result=docker(*args,timeout=45);evidence=json.loads(result.stdout);assert len(evidence['observations'])==14 and all(x['passed']for x in evidence['observations'])
finally:
remaining=docker('ps','-aq','--filter','name=^/'+name+'$').stdout.strip
if remaining:docker('rm','-f',name)
assert not docker('ps','-aq','--filter','name=^/'+name+'$').stdout.strip
report=dict(executedAt=datetime.datetime.now(datetime.timezone.utc).isoformat,files={n:hashlib.sha256((here/n).read_bytes).hexdigestfor n in ['run.py']},dockerVersion=info['ServerVersion'],imageReference=image,container=name,execution=evidence,cleanup=dict(ownedContainerRemoved=True,hostSecurityChanged=False),scope='Actual Linux aarch64 process,filesystem,credential and additional seccomp-filter observations. Built-in Docker seccomp retained. No Kubernetes,AppArmor,SELinux,host firewall,full isolation audit,performance benchmark or official practical exam execution.')
out.write_text(json.dumps(report,indent=2)+'\n');print('PASS14',evidence['python'],evidence['kernel'])
if __name__=='__main__':
if '--inside' in sys.argv:probe
else:run
A query changes from success to EPERM after an additional filter, but the flag observed through /proc remains set.
Common pitfalls
Attributing EPERM to one cause; confusing disable with stop; treating privileged mode as proof of correctness; inferring complete isolation from one local test.
Related topics: Identities and delegation · Profiles and workload admission · Incident investigation
Each control has a scope. Acceptance connects configuration to the real process and distinguishes recovered function from proven restriction.
Reference: CKS certification and domains · Kubernetes v1.35; current six-domain CKS outline