← Kubernetes: operate workloads and recover services
07 / 8 · 60 MIN

Rollout, capacity and acceptance criteria

Plan Deployment changes with calculated allowances, availability observation and recovery compatible with application state.

Identify what actually changed

Start review with the Pod-template diff. An annotation on external Deployment metadata is not equivalent to a template change, and replica scaling alone does not create a new application revision. In the fictional exercise, the ticket mixes rule updates, capacity increases and image changes. Separate each operation’s intent and identify expected evidence. Record version, scope, owner and acceptance condition. The team can then explain why more Pods appeared without attributing an application change to the controller when that change was not requested.

Calculate allowances before choosing the window

With three replicas and both percentages at 25%, maxSurge permits one additional replica and maxUnavailable resolves to zero. Calculation rounds the former up and the latter down. Do not interpret these allowances as a guarantee against every failure. In the capacity model, each new Pod requests 500m while 300m eligible capacity remains: CPU is short by 200m. The example excludes init containers, sidecars and overhead and does not resolve memory or affinity. An authorized replica allowance does not create physical capacity to run the next Pod.

Connect readiness, stabilization and progress

A Pod that has been Ready for ten seconds may not count as Available when minReadySeconds requires thirty. Observe continued readiness and relevant crashes before concluding the controller is delayed. The progress deadline must exceed that period and match service behavior. The initial lesson already separates deadline from rollback; here add the connection between timing and criteria. In the closing case, a new process answering a simple endpoint may still fail to complete the required business operation. Acceptance needs both operational and functional views of the result.

Pausing and termination change state interpretation

While a Deployment is paused, template changes do not trigger new rollouts. Review accumulated changes before resuming with appropriate authorization and observation. During replacement, old Pods may remain Terminating and consume resources. A budget ignoring this interval may fail despite apparently sufficient surge configuration. Distinguish intent, creation, availability and completed termination. Prepare handover with timestamps, current revision, pending changes and observed capacity. Do not infer resource release merely because a termination marker appeared; the previous process may still be draining work or waiting for its deadline.

Reversal needs compatible artifacts and data

The exercise sets revisionHistoryLimit=0 and establishes that old ReplicaSets were already removed. In that state, the plan should not promise recovery of that revision through rollout undo. Locate a previously approved declaration and assess how to recover it. If the release changed data into an incompatible format, reapplying the old template does not resolve that dependency. Involve application and data owners to choose between a forward fix, compatibility work or rehearsed recovery. Preserve later results and uncertain-operation identities. An available command does not replace a procedure with understood conditions and consequences.

Workshop: defend a go or no-go decision

Prepare a sheet for the fictional funds API: three replicas, 25% allowances, new Pending Pods and a functioning old service. Explain the calculation, present constraint events and propose options with explicit impact. Do not lower requests merely to meet the schedule. Include functional outcome, backlog and the decision owner. Handover can say: “The current service remains available; the new revision is waiting for eligible capacity.” This is an original documentary scenario. The following lab generates manifests locally without executing the scheduler, controller or application.

IN PRACTICE

Fictional model: three replicas, surge 25% → 1, unavailable 25% → 0. A 500m Pod with 300m eligible remaining capacity has a 200m CPU shortfall; other criteria still need assessment.

Common pitfalls

Confusing scaling with revision; rounding both percentages up; treating Ready as immediate Available; ignoring Terminating; promising reversal after history removal and data changes.

Related topics: Workloads and desired state · Resources and scheduling · Configuration and data

Take this idea with you

Rollout needs operational headroom and functional evidence. Use configuration to plan, observe actual state and maintain a recovery path compatible with data.

Create account

Reference: Deployments · Kubernetes v1.37 concepts; current official documentation consulted 2026-09-30; cluster versions and plugin capabilities must be confirmed

Kubernetes® is a registered trademark of The Linux Foundation. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by The Linux Foundation. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.