← CKS: Kubernetes security in production
20 / 20 · 120 MIN

Recover services and validate Linux controls

Distinguish modes, descriptors, threads and loaded policies to contain and recover services with evidence.

1. Define the state the change must produce

An APS team is preparing reconciliation close when it receives a hardening recommendation. The ticket asks it to remove an unnecessary daemon, restrict batch-file access and apply a syscall filter. These are three different changes with different evidence requirements. A completed command does not establish that the intended state exists. This fictional banking case does not represent internal BNP Paribas procedures. Start by identifying the component, owner, consumers, version and recovery impact. A package may have been updated while a process still uses the earlier executable. A unit may be stopped while another activates it through a socket. A file may have a different mode while a process retains an open descriptor. One thread may be filtered while another still uses the earlier policy. Each situation needs an observation at the level where its control acts. Build a table with previous state, change, positive evidence and negative evidence. Positive evidence confirms the legitimate flow, such as completing a test reconciliation. Negative evidence confirms the restriction, such as a denied open or syscall. Add boundaries: filesystem, architecture, thread, node and activation mechanism. Before promotion, ask another operator to explain what each result establishes. If they can only repeat the executed command, acceptance evidence still needs work. Assign an owner to resolve each missing observation.

2. Permissions, ownership and open references

The exercise uses only synthetic files on local tmpfs. It opens a file for reading and writing, removes access bits and tries to open it again. The fresh open receives EACCES. The earlier descriptor still reads and writes. This is not a bypass invented by the script: the two operations use different references and checks. The conclusion is limited to the filesystem and controls actually exercised. The process owns the file and has no effective, permitted or bounding capabilities. It can nevertheless restore mode 0600. That observation rules out an oversimplified compromise response: dropping capabilities does not automatically remove the owner’s authority over file mode. If an attacker controls the owning account, containment based only on chmod can be reversed by that account. Assess identity, existing processes and ownership before choosing the measure. Next, the script replaces the pathname with another file. A fresh read obtains new state; the old descriptor retains the old inode, now without links. Compare pathname stat against descriptor fstat to understand the difference. Do not assume corruption merely because readers disagree. Do not transfer the conclusion to NFS either: a remote server may check access on later operations. This lab runs no NFS, AppArmor or SELinux and establishes no behavior for those mechanisms. Document those exclusions in the acceptance record.

3. Check coverage per thread

The second part creates a worker before installing the seccomp filter. Main and worker execute the same direct getppid syscall. An additional filter returning EPERM for that syscall is then installed on the main thread with flags=0. Main starts receiving denial; the earlier worker still gets the normal result. A new thread created by main inherits its creator’s filter and receives denial. This design separates three situations that a single /proc read can hide: the thread installing the control, a thread already present and a thread created later. The script queries state by TID and runs queue-synchronized probes, avoiding attribution to one worker of a result observed on another. Docker’s builtin profile is not changed; the new filters add restrictions to existing policy. Next, an installation with TSYNC synchronizes threads in this compatible case. The worker begins receiving EPERM and both threads show two additional filters. The exercise does not construct a divergent tree that makes TSYNC fail. Questions address that case using the documentation: without the TSYNC_ESRCH variant, a positive return can identify the thread preventing synchronization. Do not use “anything other than -1 means success”. Record return values, errno when applicable and each thread’s behavior before claiming coverage. Keep the evidence tied to the exact process and policy generation tested.

4. Run in a disposable environment

Save the lesson code as run.py and run python3 run.py --output result.json on a workstation with Linux ARM64 Docker. The output file must not already exist. Image python@sha256:2d9aefe2fef018a7eb2c13064c89c71929800fd2e5dccdbf52ea5da5bb8d929a must be available locally: the script uses --pull=never. Reference runs used Docker 29.1.3, Python 3.13.16 inside the container and Linux 6.12.54-linuxkit. The script does not support x86; syscall numbers and ABI cannot be ported by simply copying them. The container has its own random name, networking disabled, UID/GID 1000, capabilities dropped, no-new-privileges and builtin seccomp. Its root is read-only, /tmp is a bounded tmpfs, and the only bind mount contains this script, also read-only. Memory, CPU and process limits apply. Do not mount production directories to adapt this exercise. Neither a privileged container nor changes to host security profiles are required. Before running it, predict results for each phase: fresh open, old descriptor, owner modification, inode replacement and thread coverage. Compare them against the JSON report afterwards. The script removes only the container it created and temporary files disappear with the exercise. If execution fails because of architecture, runtime or syscall behavior, retain the error for diagnosis and preserve the restriction that caused it; do not remove controls merely to obtain PASS.

5. Read results and choose recovery

Both final runs recorded 24 checks with the same code hash. They cover effective process state, absence of capabilities, reading and writing through the old descriptor, denial of a fresh open and adoption of the new inode only through pathname reading. The second half demonstrates thread-local filtering, inheritance by a new thread and successful TSYNC for the existing worker. A different syscall continues to work, providing a positive control. These checks do not measure performance or validate a real banking application. Getppid was chosen as an observable synthetic operation without executing an exploit. One operation succeeding does not establish that the whole profile is correct; an application may use other syscalls only during file close, recovery or load. Build functional tests from the concrete flow and include uncommon paths that the service must retain. If an installed filter blocks a necessary operation, adding ALLOW does not remove a denial from an earlier filter. Recovery may require a replacement instance created with the profile approved through its runtime or supervisor. An exec within the already filtered process is not, by itself, that recovery, because filters can be inherited. Define where the clean instance originates, how it takes over work and how its effective policy is confirmed as approved. Do not describe a new PID alone as proof that earlier restrictions disappeared.

6. Reduce services while retaining the management path

When retiring a diagnostic daemon, identify sockets, timers and other mechanisms that may activate it. A stopped service unit does not prove its listener has disappeared. The systemd documentation distinguishes a service from the socket unit that receives traffic and may activate it. The right action depends on dependencies and authorized function: do not switch off every socket on a host merely to simplify inventory. If you change a unit, distinguish on-disk configuration, configuration known to the manager and process state. Daemon-reload reloads manager configuration; it is not equivalent to restarting the service or automatically applying every effect to an already running process. Similarly, a new package may leave a process associated with its earlier executable, visible through /proc/PID/exe. Define the observation that establishes the new effective generation. For the batch, replacement must respect checkpoints, dependencies and duplicate prevention. Stopping a process without confirming its last reconciled item may create a functional incident while you address the security incident. Coordinate the window with APS and business, retain a tested management path and define who can authorize recovery. This lesson runs neither systemd nor package updates: these are documented scenarios applying the effective-state reasoning observed in the Linux exercise. Their production procedures require validation in the actual service environment.

7. Validate network reach and profile adoption

Test an SSH change with identified source, user, destination address and network family. An IPv4 denial does not establish IPv6 denial if both families remain configured. With Match rules, syntax validation alone does not explain the configuration selected for a connection. The sshd -T mode with -C parameters helps inspect that selection, but the effective daemon still needs testing and an approved recovery access path must remain available. A protocol option also has limited scope. Disabling AllowTcpForwarding restricts forwarding provided by SSH; it does not automatically turn an allowed shell into an environment without forwarding possibilities. If isolation is required, assess available tools, reachable destinations and additional controls. Do not confuse authentication rejected for a missing key with the network denial you intended to demonstrate. For AppArmor, distinguish distributed file, loaded profile, mode and process association. Documentation explains that file changes need a reload to take effect. A profile in complain does not establish that the observed operation was prevented; the negative test must show the intended result. Cover nodes where the workload can run. This module executed no SSH, AppArmor, firewall or Kubernetes. Questions about those mechanisms are original applications of sources, with explicit boundaries around the practice actually performed. Record those boundaries in the handover so another team does not mistake scenarios for executed acceptance.

8. Hand decisions and evidence to RUN

In the final case, the batch main thread is filtered while an earlier worker continues executing the syscall. The first decision is to recognize uneven coverage, preserve necessary state and measure relevant threads. If TSYNC is the supported solution, confirm its return value and results. A synchronization failure cannot become success merely because code checks only for -1. If the change breaks the legitimate flow, prepare the approved recovery instance. The technical manager coordinates whoever controls the process, whoever validates policy and whoever confirms functional integrity. Define the decision point before closing and the impact of waiting, replacing or activating contingency. Use checkpoints to avoid repeating movements; do not promise continuity simply because a new process started. Business must confirm reconciliation results, while APS confirms observability, access and operational support. The handover includes versions, code or profile hash, effective identity, thread matrix, file references, expected and observed results, cleanup and limitations. Keep open questions separate from completed evidence. Summarize the related concepts: object references, control composition, inheritance, reconciliation and safe change. This path’s internal assessment trains decisions; it does not replace the official practical exam or establish complete competence merely through a score. Independent specialist review remains necessary. The next shift should be able to repeat the bounded acceptance checks and identify exactly which claims still need environment-specific evidence.

"""Original practice: local file references and seccomp thread synchronization.

Host entrypoint creates only an owned, restricted, disposable Linux container.
"""
import sys


def probe:
 import ctypes
 import errno
 import json
 import os
 import platform
 import queue
 import tempfile
 import threading
 from pathlib import Path

 assert platform.system == 'Linux' and platform.machine == 'aarch64'
 libc = ctypes.CDLL(None, use_errno=True)
 libc.syscall.restype = ctypes.c_long
 records = []

 def check(name, observed, expected):
 assert observed == expected, (name, observed, expected)
 records.append(dict(name=name, observed=observed, expected=expected, passed=True))

 def status(tid):
 raw = dict(line.split(':', 1) for line in Path('/proc/self/task/' + str(tid) + '/status').read_text.splitlines if ':' in line)
 return {k: raw[k].strip for k in ['Uid', 'Gid', 'CapEff', 'CapPrm', 'CapBnd', 'NoNewPrivs', 'Seccomp', 'Seccomp_filters']}

 def syscall(number):
 ctypes.set_errno(0)
 result = libc.syscall(ctypes.c_long(number))
 return dict(result=result, errno=ctypes.get_errno)

 main_tid = threading.get_native_id
 initial = status(main_tid)
 check('restricted-process', [os.getuid, os.getgid, initial['NoNewPrivs'], initial['Seccomp']], [1000, 1000, '1', '2'])
 check('capability-sets-empty', all(int(initial[k], 16) == 0 for k in ['CapEff', 'CapPrm', 'CapBnd']), True)
 # Local tmpfs only; do not infer NFS or mandatory-access-control behavior.
 with tempfile.TemporaryDirectory(prefix='dr-system-response-') as directory:
 target = Path(directory) / 'state'
 target.write_bytes(b'old-state')
 fd = os.open(target, os.O_RDWR)
 try:
 check('initial-descriptor-read', os.pread(fd, 9, 0).decode, 'old-state')
 target.chmod(0)
 try:
 new_fd = os.open(target, os.O_RDONLY)
 os.close(new_fd)
 open_errno = 0
 except OSError as error:
 open_errno = error.errno
 check('new-open-denied-after-chmod', open_errno, errno.EACCES)
 check('existing-descriptor-still-reads', os.pread(fd, 9, 0).decode, 'old-state')
 os.pwrite(fd, b'old-write', 0)
 check('existing-descriptor-still-writes', os.pread(fd, 9, 0).decode, 'old-write')
 target.chmod(0o600)
 check('owner-restores-mode-without-capabilities', target.stat.st_mode & 0o777, 0o600)
 check('fresh-open-after-owner-restores-mode', target.read_text, 'old-write')
 replacement = Path(directory) / 'replacement'
 replacement.write_bytes(b'new-state')
 os.replace(replacement, target)
 check('replacement-changes-path-inode', os.fstat(fd).st_ino!= target.stat.st_ino, True)
 check('old-descriptor-retains-old-inode-data', os.pread(fd, 9, 0).decode, 'old-write')
 check('fresh-path-read-uses-replacement', target.read_text, 'new-state')
 check('old-inode-unlinked-but-open', os.fstat(fd).st_nlink, 0)
 finally:
 os.close(fd)

 check('synthetic-files-removed', not Path(directory).exists, True)
 # Native aarch64 ABI only: getpid=172, getppid=173, seccomp=277.
 baseline_parent = syscall(173)
 check('main-getppid-initially-allowed', baseline_parent['result'] >= 0 and baseline_parent['errno'] == 0, True)
 inbox, replies = queue.Queue, queue.Queue

 def worker:
 replies.put(dict(tid=threading.get_native_id, query=syscall(173)))
 while True:
 message = inbox.get
 if message == 'stop':
 return
 replies.put(dict(query=syscall(173), status=status(threading.get_native_id)))

 thread = threading.Thread(target=worker)
 thread.start
 worker_start = replies.get(timeout=5)
 check('existing-worker-initially-allowed', worker_start['query'], baseline_parent)
 worker_initial = status(worker_start['tid'])

 class Filter(ctypes.Structure):
 _fields_ = [('code', ctypes.c_ushort), ('jt', ctypes.c_ubyte), ('jf', ctypes.c_ubyte), ('k', ctypes.c_uint)]

 class Program(ctypes.Structure):
 _fields_ = [('len', ctypes.c_ushort), ('filter', ctypes.POINTER(Filter))]

 def install(flags):
 # Reject unexpected ABI, deny getppid, otherwise retain preceding filters.
 rows = [(0x20, 0, 0, 4), (0x15, 1, 0, 0xc00000b7), (0x06, 0, 0, 0x80000000),
 (0x20, 0, 0, 0), (0x15, 0, 1, 173), (0x06, 0, 0, 0x00050000 | errno.EPERM),
 (0x06, 0, 0, 0x7fff0000)]
 filters = (Filter * len(rows))(*(Filter(*row) for row in rows))
 program = Program(len(rows), filters)
 ctypes.set_errno(0)
 result = libc.syscall(ctypes.c_long(277), ctypes.c_uint(1), ctypes.c_uint(flags), ctypes.byref(program))
 return dict(result=result, errno=ctypes.get_errno)

 denied = dict(result=-1, errno=errno.EPERM)
 try:
 check('thread-local-filter-installed', install(0), dict(result=0, errno=0))
 check('main-getppid-now-denied', syscall(173), denied)
 inbox.put('probe')
 still_allowed = replies.get(timeout=5)
 check('existing-worker-not-synchronized-by-flags-zero', still_allowed['query'], baseline_parent)
 inherited = queue.Queue
 child_thread = threading.Thread(target=lambda: inherited.put(syscall(173)))
 child_thread.start
 child_thread.join(timeout=5)
 assert not child_thread.is_alive
 check('new-thread-inherits-creator-filter', inherited.get(timeout=5), denied)
 # The worker tree remains an ancestor of the caller tree, so TSYNC can succeed.
 check('tsync-filter-installed', install(1), dict(result=0, errno=0))
 inbox.put('probe')
 synchronized = replies.get(timeout=5)
 check('existing-worker-denied-after-tsync', synchronized['query'], denied)
 after = status(main_tid)
 check('both-threads-have-two-additional-filters', [int(after['Seccomp_filters']) - int(initial['Seccomp_filters']),
 int(synchronized['status']['Seccomp_filters']) - int(worker_initial['Seccomp_filters'])], [2, 2])
 check('unrelated-getpid-still-allowed', syscall(172)['result'] == os.getpid, True)
 finally:
 inbox.put('stop')
 thread.join(timeout=5)
 assert not thread.is_alive
 check('worker-thread-ended', thread.is_alive, False)
 print(json.dumps(dict(python=platform.python_version,kernel=platform.release,architecture=platform.machine,
 observations=records,initialStatus=initial,workerInitialStatus=worker_initial)))


def run:
 import argparse
 import datetime
 import hashlib
 import json
 import subprocess
 import uuid
 from pathlib import Path
 parser = argparse.ArgumentParser
 parser.add_argument('--output', required=True)
 args = parser.parse_args
 out = Path(args.output)
 assert not out.exists
 image = 'python@sha256:2d9aefe2fef018a7eb2c13064c89c71929800fd2e5dccdbf52ea5da5bb8d929a'
 name = 'dr-cks-system-response-' + uuid.uuid4.hex[:12]

 def docker(*arguments, timeout=30):
 result = subprocess.run(['docker', *arguments], capture_output=True, text=True, timeout=timeout)
 if result.returncode:
 raise RuntimeError(result.stderr[:1500])
 return result.stdout

 info = json.loads(docker('info', '--format', '{{json.}}'))
 assert info['Architecture'] == 'aarch64'
 assert any('seccomp' in option for option in info['SecurityOptions'])
 try:
 result = docker('run','--rm','--pull=never','--name',name,'--label','dr.lesson=cks-system-response',
 '--network=none','--read-only','--user','1000:1000','--cap-drop=ALL',
 '--security-opt=no-new-privileges=true','--security-opt=seccomp=builtin',
 '--pids-limit=32','--memory=64m','--cpus=0.5',
 '--tmpfs','/tmp:rw,noexec,nosuid,nodev,size=1m,mode=1777',
 '--mount','type=bind,src='+str(Path(__file__).resolve)+',dst=/lesson/run.py,readonly',
 '--entrypoint','python',image,'-B','/lesson/run.py','--inside',timeout=45)
 evidence = json.loads(result)
 assert all(row['passed'] for row in evidence['observations'])
 finally:
 if docker('ps','-aq','--filter','name=^/'+name+'$').strip:
 docker('rm','-f',name)
 assert not docker('ps','-aq','--filter','name=^/'+name+'$').strip
 report = dict(executedAt=datetime.datetime.now(datetime.timezone.utc).isoformat,
 scriptSha256=hashlib.sha256(Path(__file__).read_bytes).hexdigest,
 dockerVersion=info['ServerVersion'],imageReference=image,container=name,passed=True,execution=evidence,
 cleanup=dict(ownedContainerRemoved=True,hostSecurityChanged=False),
 scope='Actual local tmpfs DAC and descriptor observations, aarch64 thread-local seccomp and successful TSYNC on compatible filter trees. No NFS, AppArmor, SELinux, systemd, SSH, firewall, Kubernetes, divergent-tree TSYNC failure, production recovery or full practical mock executed.')
 out.write_text(json.dumps(report,indent=2)+'\n')
 print('PASS',len(evidence['observations']),evidence['python'],evidence['kernel'])


if __name__ == '__main__':
 if '--inside' in sys.argv:
 probe
 else:
 run
IN PRACTICE

Main receives EPERM while an earlier worker still executes getppid until successful TSYNC synchronization.

Common pitfalls

Confusing mode with descriptor revocation, a file with loaded policy, the main thread with every worker and a new PID with clean policy.

Related topics: Linux permissions and ownership · Seccomp and TSYNC · AppArmor and profile adoption · Recovery and handover

Take this idea with you

Acceptance must observe the object, reference and thread where the control actually acts.

Create account

Reference: CKS domains and exam details · Kubernetes v1.35; current six-domain CKS outline

Kubernetes® and CKS are trademarks or registered trademarks of The Linux Foundation. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by The Linux Foundation. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.