Follow the request from name to result
A useful infrastructure review starts with a concrete business request. In our fictional example, operations submits a fund close and expects a result within fifteen minutes. Draw the path: queried name, resolver, network connection, backend, storage and external dependencies. For every step, identify available evidence and what remains unproven. An available tunnel does not prove DNS resolution; a healthy backend does not prove session preservation; a completed task does not prove the whole document was processed. This decomposition helps APS coordinate networking, systems, middleware and development without passing the incident indefinitely between teams. Start with a trial carrying a correlation identifier and timestamps, preserving its source context. Then repeat through an anticipated degraded path. The review question is always which requirement survives that state. Do not add a resilience promise based on an indicator describing only part of the path.
Distinguish resolution, routing and NAT resources
Two overlapping private zones can explain NXDOMAIN even when one contains a record. For gateway.settlement.corp.example, authorized zone settlement.corp.example is more specific than corp.example. If its record is missing, repeating the query does not automatically search the less specific zone. Record the zone actually selected before changing routes. For on-premises forwarding, also observe the packet’s real source: return routing for 35.199.192.0/19 must pass through the VPC that originated the query. A route only to the VM subnet does not cover that requirement. A Cloud NAT OUT_OF_RESOURCES error instead directs investigation toward translation resources, even with low bits per second. Examine per-VM usage, concurrency and connection reuse. These are three different investigations with different evidence. In an incident meeting, request an observation that distinguishes hypotheses and a bounded change that can be validated. Avoid broad permission changes or increased retries before identifying the failing mechanism.
Count only capacity that can serve
Usable capacity depends on more than the sum of free resources. A Pod may need memory, an appropriate zone and mandatory spread simultaneously. If free nodes violate DoNotSchedule and the others lack memory, aggregate capacity exists but a valid placement does not. Changing the rule to ScheduleAnyway changes the requirement; it may be an approved mitigation but should not be presented as proof of the original design. Also count physical resources correctly in the budget: three nodes per zone across three zones means nine node VMs. The managed control plane is not another part of that pool. On the traffic side, session affinity can break when the backend becomes unavailable. If continuity depends on state, ask where its recoverable copy lives and how another backend obtains it. Use the exercise to present effective capacity and its calculation conditions to the sponsor, rather than just a desired replica count.
Read requests, metrics and limits separately
This lesson’s YAML excerpt contains only container resources and an HPA metric; it is not a manifest to apply. With four ready single-container Pods, each requesting 500m and averaging 300m usage, relevant utilization is 60% of request. Against a 50% target, the base formula gives ceil(4 × 60/50), or five replicas. Before calculating, identify the denominator: using a full core produces a different and incorrect conclusion for this configuration. This calculation does not simulate the controller. Missing metrics, tolerance, bounds and policies can affect the real action; investigate them in the cluster. Memory has a different interpretation: a container can terminate under pressure at its limit even when the node has free memory. Correlate OOMKilled events with cgroup observations and consumption profiles. Requesting more CPU or changing readiness does not automatically address that cause. The exercise aims to select relevant evidence for each mechanism rather than memorize a replica count.
Validate storage from guest to recovery
Enlarging a disk resource does not demonstrate that the application gained usable space. For an expanded data Persistent Disk, check device, partition, filesystem and mount point. Expansion procedure depends on layout; do not format a filesystem containing data to resolve a missing growth operation. Performance also has boundaries: when throughput is limited by the VM, doubling a disk size already above that limit does not double effective capacity. Measure the complete path with a representative workload. For recovery, independent snapshots of data and journal during writes do not prove one logical state. Define application coordination and demonstrate joint recovery. Finally, a Cloud Storage FUSE mount should not be accepted merely because the application can list files. If it depends on file locking among writers, that need must guide selection. In each case, the learner should write an acceptance condition and a trial capable of disproving a technically plausible but incomplete proposal.
Budget execution time and memory
A Cloud Run task timeout applies to each attempt. A job with a ten-minute timeout and two retries does not demonstrate a ten-minute global deadline, or even a fifteen-minute deadline. Record attempt durations, intervals and when the useful result becomes available. If there is a business window, define an appropriate control for that window and decide what happens to partial effects before retrying. For interactive services, include temporary files in the memory budget: the container’s writable filesystem uses instance memory. A stable heap can coexist with growing consumption from files retained between requests. Compare streaming, controlled cleanup and bounded accumulation against real export formats and sizes. More resources may be necessary, but should come with an explanation of growth and load conditions. RUN handover should include signals distinguishing timeout, memory pressure and dependency failure, with actions maintaining traceability of the business result.
Verify the meaning of AI outputs
Inference endpoint availability does not prove that input has the meaning used during training. In the original example, training transforms x into (x-10)/2 while serving uses x/100. For 14, the model receives 2 through one path and 0.14 through the other. Align and version preprocessing in the Predictor and trial known examples through the real path before attributing the difference to model quality. In Speech-to-Text, boosting terms can reduce omissions while increasing false positives. The evaluation set needs calls both with and without those terms, representing the operational impact of both errors. With asynchronous OCR, PDF results can occupy several files: the consumer must reconcile pages and response errors instead of assuming the first object represents the whole document. These validations are part of the useful infrastructure supporting a process. Do not turn an HTTP response, transcription or extracted text into automatic authorization for a sensitive operation.
Prepare a defensible acceptance decision
The final case combines a name that fails to resolve for the application and a batch exceeding its window despite green partial indicators. Prepare a short committee note: requirement, observation, consequence, action and owner. The DNS correction must consider other zone names and consumers; rushed deletion can move the problem elsewhere. The timing trial must include retries and useful output, not just one attempt’s duration. If the sponsor changes the requirement, record an explicit decision and accepted risks without retrospectively rewriting the trial as successful. This lesson’s technical summary is to follow the real path, count eligible capacity, distinguish requests from limits, check storage semantics and measure functional completion. Connect these subjects to observability, change management and handover criteria. As a communication exercise, explain the decision in English to an international team without relying on vague terms such as cloud ready. The listener should be able to identify the next trial and who supplies the evidence.
# Guided excerpts only; not a deployable manifest.
# One container per Pod in the numerical exercise.
resources:
requests:
cpu: 500m
memory: 256Mi
limits:
memory: 512Mi
---
# HPA metric fragment; four ready Pods averaging 300m each.
type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 50
An application cannot resolve its endpoint and a batch exceeds its window: network and isolated-attempt indicators do not establish useful delivery.
Common pitfalls
Confusing desired replicas with eligible capacity, disk space with filesystem space, per-attempt timeout with a global deadline and partial AI results with complete coverage.
Related topics: Workload capacity and availability · Observability and functional recovery · Migration and acceptance criteria
Validate the complete path and result meaning; size and accept each resource according to the boundary it actually controls.
Reference: Cloud DNS zones overview · Current linked standard guide; edition date unconfirmed (2026-09-30 inspection)