Define what each signal establishes
An operator needs to distinguish observation, hypothesis and conclusion before choosing an action. A closed alert, a Ready node and an HTTP 200 endpoint describe different facts. None alone proves that a reconciliation process is producing the correct result again. Start investigation with consumer impact and build an evidence chain: request, response, processing, produced data and confirmation by the next system. Record identity, version, time interval and observed population to avoid comparing signals from different executions. In the exercise, receive a fictional dossier with four green indicators and one failed functional test. For every indicator, write the limited claim it supports and the verification still missing. Mitigation priority should follow impact and the authorized procedure. The objective is to let another person continue the investigation without turning an initial hypothesis into a supposedly confirmed cause.
Preserve resets and distributions in metrics
Counters from different instances can reset at different times. If they are summed before calculating rate, one instance’s growth can hide another’s decrease. Calculate rates per original series before aggregation, preserving labels needed for identity. Also avoid combining percentiles as if they were means. Two local p95 values, even with equal counts, do not describe the tails determining combined p95. This lesson’s Python exercise uses nearest-rank percentiles defined over integer observations without interpolation. In two datasets, local counts are 100 and p95 values are 10 and 100 ms; global p95 changes from 100 to 50 ms while the local average remains 55. Run the example and explain the lost information. In real observability, choose aggregatable distributions and appropriate resolution, documenting approximations. The exercise does not simulate the Prometheus engine or validate a vendor’s histogram interpolation.
Review the population selected by a query
A filter can be syntactically valid while excluding events needed for diagnosis. In Cloud Logging, comparison against a missing jsonPayload field silently fails. In the original example, A has stage ready, B has stage failed and C has no stage. The stage!=ready comparison selects B but not C. Negating equality includes B and C because equality is true for neither. Use the full field path in the actual query and distinguish missing fields from defined fields with default values. Before using counts to accept a release, trial representative entries including absence and unexpected values. The exercise asks you to classify each event by hand and explain whether the intent is to select only known failures or also unknown states. Then document that intent in the runbook. A query omitting unknown states does not establish that every excluded event succeeded.
Separate application, Pod and node health
Running is not equivalent to Ready. An application can keep its process active while a dependency prevents serving requests. Readiness lets a Pod leave normal Service traffic without restarting the container merely for that failure. Liveness and startup probes have other purposes. During updates, a PDB does not replace Deployment strategy: workload controllers are not constrained by the PDB during application rolling updates. Review workload settings and effective capacity before promising minimum availability. In GKE, node auto-repair responds to node-health conditions; it does not fix invalid application configuration on healthy nodes. The exercise shows CrashLoopBackOff in a new version with Ready nodes. Identify the failing layer, authorized mitigation and signals that would prove recovery. Recreating infrastructure or removing readiness without fixing the cause can prolong unavailability and make diagnosis harder for the next shift.
Validate the response and alert lifecycle
An endpoint can return HTTP 200 with a login page instead of the expected document. Configure content validation appropriate to the contract and consumer access context. A connectivity test remains useful, but identify its scope and complement it with functional verification. The example JSON is small; it does not promise unlimited inspection of every response by uptime checks. Also distinguish alert lifecycle from service health. If autoclose closes an alert after telemetry stops arriving, closure does not prove the metric returned to normal. Investigate lost observation and run an independent test before declaring recovery. In the exercise, write two committee messages: one describes only alert state; the other describes known service state and the next verification. This separation prevents administrative closure from prematurely ending operational investigation.
Reconcile messages and recovery points
Pub/Sub exactly-once has conditions and does not remove the need for correct acknowledgment handling. In this lesson’s case, a pull subscription operates in one region, but its ack ID expired and is rejected with INVALID_ARGUMENT. Retrying that ID indefinitely does not acknowledge the current delivery. Manage deadlines and reconcile effects already produced before repeating business work. A dead-letter policy’s maximum attempt count is approximate even with correct IAM; do not use it as an exact execution counter. For Cloud SQL recovery, inspect the effective PITR window and compare it with the requested instant. Fourteen days of backups do not establish that every instant in those days has the logs required for PITR. The exercise asks for an evidence inventory: delivery identifier, effect state, recoverable window and decision on acceptable loss. Evaluate each mechanism against its actual guarantee while keeping the business objective visible in the technical decision.
Turn incidents into verifiable improvements
A blameless postmortem retains the factual sequence and examines conditions that made error possible. If an operator followed an approved runbook lacking target validation, and two environments had almost identical names, repeating only the person’s name does not fix the risk. Define actions improving identification, validation and trials, with an owner and completion evidence. Release governance also needs explicit rules. An error-budget policy can suspend features while permitting urgent security fixes under separate review. Applying that exception does not mean relabeling a feature to bypass the limit. In the exercise, analyze the supplied fictional policy and classify two changes, explaining each approval path. Avoid turning a vendor’s example policy into a universal rule for every organization. Value comes from consistent, previously agreed decisions linking reliability and risk to work the team performs after the incident.
Prepare acceptance and continuity across shifts
In the final cases, explain why the service cannot yet be accepted from the presented evidence. One case combines incorrect content with a query excluding relevant events. The other combines an unavailable rollout, invalid configuration and excessive reliance on PDB and auto-repair. Produce a short decision with impact, containment, owner and next acceptance condition. Link every action to the correct mechanism and avoid declaring recovery merely because a resource returned to an expected administrative state. Keep local exercise results and explain the percentile definition so another colleague can reproduce the calculation. During international handover, distinguish confirmed facts, open hypotheses and ongoing operations, including effects still needing reconciliation. These exercises are original and fictional. Python validates finite local data; it executes neither Kubernetes, Prometheus, Pub/Sub nor Cloud SQL and does not replace authorized trials in the actual environment.
"""Original finite-data exercise; no Prometheus or cloud service is executed."""
from math import ceil
from pathlib import Path
import hashlib
import json
import platform
def nearest_rank(values, percentile):
if not values or not isinstance(percentile, int) or not 1 <= percentile <= 100:
raise ValueError('Nonempty observations and an integer percentile 1..100 are required')
ordered = sorted(values)
return ordered[ceil(percentile * len(ordered) / 100) - 1]
def analyze(tail):
a = [10] * 95 + [tail] * 5
b = [20] * 94 + [100] * 6
local = [nearest_rank(a, 95), nearest_rank(b, 95)]
return {'counts': [len(a), len(b)], 'local_p95_ms': local,
'mean_local_p95_ms': sum(local) / 2,
'global_p95_ms': nearest_rank(a + b, 95)}
def main:
checks = []
def check(name, condition):
if not condition:
raise AssertionError(name)
checks.append(name)
high_tail = analyze(1000)
lower_tail = analyze(50)
check('both fixtures retain equal counts', high_tail['counts'] == lower_tail['counts'] == [100, 100])
check('both fixtures retain identical local percentiles', high_tail['local_p95_ms'] == lower_tail['local_p95_ms'] == [10, 100])
check('both averages are 55', high_tail['mean_local_p95_ms'] == lower_tail['mean_local_p95_ms'] == 55)
check('high tail gives global 100', high_tail['global_p95_ms'] == 100)
check('lower tail gives global 50', lower_tail['global_p95_ms'] == 50)
check('same retained summaries cannot determine global percentile', high_tail['global_p95_ms']!= lower_tail['global_p95_ms'])
check('mean of local percentiles fails in both fixtures', all(x['global_p95_ms']!= x['mean_local_p95_ms'] for x in [high_tail, lower_tail]))
boundary_cases = 0
for n in range(1, 41):
values = list(range(1, n + 1))
for percentile in [50, 90, 95, 99, 100]:
expected_rank = (percentile * n + 99) // 100
if nearest_rank(values, percentile)!= expected_rank:
raise AssertionError((n, percentile))
if nearest_rank(list(reversed(values)), percentile)!= expected_rank:
raise AssertionError(('order', n, percentile))
boundary_cases += 1
check('rank and ordering checks cover 200 finite cases', boundary_cases == 200)
outcomes = set
for tail in range(11, 111):
result = analyze(tail)
if result['counts']!= [100, 100] or result['local_p95_ms']!= [10, 100]:
raise AssertionError(('summary', tail))
if result['global_p95_ms']!= max(20, min(tail, 100)):
raise AssertionError(('global', tail))
outcomes.add(result['global_p95_ms'])
check('100 tail variations preserve the retained summaries', len(range(11, 111)) == 100)
check('identical summaries allow 81 different global p95 values', len(outcomes) == 81)
for values, percentile in [([], 95), ([1], 0), ([1], 101), ([1], 95.5)]:
try:
nearest_rank(values, percentile)
except ValueError:
continue
raise AssertionError('invalid input was accepted')
check('invalid percentile inputs are rejected', True)
check('one observation has the same p95', nearest_rank([7], 95) == 7)
print(json.dumps({'scriptSha256': hashlib.sha256(Path(__file__).read_bytes).hexdigest,
'checks': len(checks), 'checkNames': checks,
'fixtures': {'highTail': high_tail, 'lowerTail': lower_tail},
'rankCases': boundary_cases, 'tailCases': 100, 'distinctGlobalResults': len(outcomes),
'python': platform.python_version, 'vendorExecution': False,
'network': False, 'persistentWrites': False, 'independentVerification': False}, indent=2))
if __name__ == '__main__':
main
Two populations with identical local counts and p95 can have different global p95 values; a green HTTP check can also represent the wrong page.
Common pitfalls
Averaging percentiles, omitting missing fields, attributing universal availability control to PDB and confusing alert closure or a Ready node with recovery.
Related topics: Observability and indicator quality · Incident management and operational learning · Resilience and service recovery
Preserve the information needed to interpret each signal and demonstrate recovery on the consumer’s required functional path.
Reference: Prometheus query functions · Current linked standard guide; edition date unconfirmed (2026-09-30 inspection)