1. Define what approval covers
An infrastructure change needs to bind intent to the objects it will alter. In an APS exercise, the ticket approves plan A for preview-west, but the runner receives plan B for prod-west. The job is still named deployment-approved. That name does not establish that approval covers the submitted package or environment. Identify configuration revision, inputs, plan, target, backend and execution identity. Keep the approval reference somewhere controlled by the agreed process. If the sender can replace both plan and approval without control, comparing hashes merely confirms two pieces the sender supplied. A checksum recognizes a specific representation. It does not prove that an authorized person reviewed actions, that quota is available or that the application tolerates the change. The project team should define entry criteria for the window and ownership of each check. Support needs to explain why execution stopped when package identity diverged. A documented mismatch is work to resolve; it should not disappear because someone copied the new hash into a ticket without reviewing content. Record the unresolved question and the person responsible for answering it.
2. Distinguish reviewed and executed plans
A plan without a saved file supports reviewing a prediction. A later apply without that file plans again. Even with unchanged code, external interventions can change the result. In a process requiring approval of concrete actions, save the plan, review it and retain the link to the artifact used for execution. Passing a saved plan to apply executes its decisions without another interactive confirmation. Job authorization therefore needs to exist before that call; an extra CLI question must not be assumed as the final control. If intent changes, create a new plan and perform the corresponding review. New planning options, such as different variable values, cannot be introduced when applying a saved plan. Editing tfvars after producing the file does not change decisions already saved in it either. During handover, distinguish “the commit was approved” from “these actions on this target were reviewed.” The first may be a necessary process condition but does not establish the second. If the window ends before new review, this exercise calls for rescheduling or the decision defined by change procedure, rather than silently reusing earlier approval.
3. Retain visibility into drift and secrets
After an intervention increases capacity from four to eight units, a plan without refresh may still reason from earlier information. The refresh=false option reduces reads but can also hide the external change. Do not use that result as proof of no drift. Obtain current observation and decide whether configuration should incorporate the change or the resource should return to prior intent. Refresh-only updates the recorded observation; it does not establish that all infrastructure matches intended configuration. Every report claim should retain that scope. Also limit conclusions after using target for exceptional recovery. Success on a targeted resource does not prove convergence elsewhere. Produce a full plan to assess what remained outside the intervention. When retaining evidence, remember that sensitive does not mean every JSON output was redacted. A show -json report can contain plaintext values. Distribute a view suited to its audience and protect the complete artifact when needed. Governance may need actions and identifiers without needing credentials. Traceability does not require publishing secrets in a ticket or broadly accessible log. Review both the artifact's contents and the permissions of its destination.
4. Recognize state and control dependencies
Lineage helps distinguish states; serial represents revisions within that identity. During state push, a different lineage or higher remote serial triggers guards. A local snapshot with serial 90 gains no legitimacy over a target with another lineage merely by having a larger number. Likewise, a file named final-approved with serial 41 may omit changes present in remote revision 44. Before a recovery write, confirm origin, destination, backups and intervening changes. Forcing a write removes guards; it does not reconcile data. Editing numbers or identities to remove an error leaves its cause unresolved. The.terraform.lock.hcl file addresses another dimension: provider versions and checksums. It does not pin remote-module versions. Control those versions through appropriate configuration and include changes in review. Running init -upgrade may select a newer provider allowed by constraints, changing the basis used to produce a plan. A mirror checksum failure requires investigating package, origin and platform. Deleting the lock file to accept whatever arrives replaces the reference without explaining mismatch. Handover should distinguish state locking, dependency selection and artifact identity; they are different checks.
5. Refactor and transfer management without duplication
When a resource changes address within configuration, a textual change alone can look to the planner like removal followed by creation. If intent is to retain the existing object, a moved block represents the association between old and new addresses in compatible Terraform versions. Review the resulting plan: the block addresses the address change but does not guarantee that unrelated attribute changes are harmless. In this exercise, type and attributes stay unchanged and the new address is free. The learner can therefore isolate the cause of an unintended destroy/create proposal. Transfer to another tool has a different intent. In a removed block, destroy=false removes the state binding without destroying the remote resource. It does not automatically install a new manager, stop costs or transfer human responsibility. Define who will maintain the resource and when the earlier pipeline stops issuing changes. Importing one object into two active states does not create coordinated management; it can create two writers with conflicting intent. Rehearse handover, retain resource-identity evidence and confirm the new plan proposes no unexpected replacement. Decommission decisions must distinguish ending management in one tool from actually ending the service.
6. Prepare upgrades with actual capacity
A GKE maintenance exclusion controls certain events but does not guarantee infrastructure will remain unchanged. Underlying services, repairs and certain critical situations can fall outside that control. The application needs to tolerate events relevant to its design. Before a planned upgrade, confirm versions, capacity, placement and dependencies as well as the calendar. The PM should explain the difference between an authorized window and technical capacity to perform the change within agreed impact. On a Standard pool with maxSurge=2 and maxUnavailable=0, strategy intends to create temporary capacity before removing existing capacity. Without extra resources, upgrade cannot progress under those conditions. Increasing maxSurge creates neither quota nor physical availability. Obtaining compatible capacity, using an appropriate reservation or reviewing strategy with explicit unavailability risk may be necessary. If some nodes were upgraded and others failed, do not assume automatic rollback of the whole pool. Observe per-node versions and record partial state. A healthy application dashboard does not establish version uniformity, just as desired version does not prove every node already runs it. Assign ownership for following the pool until the agreed end state is observed.
7. Respect the disruption budget
An application expects three replicas but only two are healthy. Its PDB requires minAvailable:2 and selects only those replicas. Removing another healthy replica through the Eviction API would leave one, so that new eviction should be blocked while conditions remain unchanged. The already unavailable replica counts against the budget. A PDB does not create a fourth replica, reserve a node or fix readiness failure. First investigate why the replica is not recovering: resources, placement, dependencies or startup itself may matter. Directly deleting a Pod to bypass a drain can bypass PDB protection. Having a budget does not turn that action into a safe disruption. The team must decide how to restore healthy headroom or obtain a different continuity plan. Also avoid promising that a PDB prevents hardware failures: this mechanism does not block involuntary disruptions. In handover, connect unavailability to the service and acceptance criterion rather than only reporting that a command waited. “Two of three replicas healthy; minimum two; no confirmed additional capacity” gives the next team a concrete decision basis. Recheck the observation before acting because readiness can change during the incident.
8. Matching and decision exercise
The Python exercise uses a small fictional report: target, lineage, serial, unknown effects and actions. This is not Terraform JSON, and the program neither interprets providers nor runs plan or apply. It compares the report with an approval record supplied to the model and evaluates six conditions: digest, target, lineage, serial, known effects and destruction scope. readyForReview only means these exercise conditions were satisfied. executionAuthorized remains false. An unauthenticated reference, a system changed after observation or a missing lock is not resolved by field equality. Run the eight examples and first predict the rejection reason. A replace without destruction approval fails even when its digest matches. A report with unknownEffects also fails under the teaching rule requiring known effects before this operational review. That rule does not assert that every unknown value in an actual Terraform plan is invalid. The 64 combinations check six conditions without external services. Finish with a short note: evidence accepted for review, conditions still to establish, target and owner. Connect that note to a real plan only through a process validating origin, freshness, access and execution. Keep the model's assumptions visible when explaining the result to another team.
"""Original evidence worksheet. Input is NOT Terraform plan JSON.
No provider calls, state locks, signatures or execution authorization. Equality
with a trusted approval record is only a classroom prerequisite for review.
"""
from copy import deepcopy
from hashlib import sha256
from itertools import product
import json
from pathlib import Path
def digest(report):
return sha256(json.dumps(report, sort_keys=True, separators=(',', ':')).encode).hexdigest
def validate(report, approval):
required = {'target', 'lineage', 'serial', 'unknownEffects', 'actions'}
if not isinstance(report, dict) or set(report)!= required:
raise ValueError('invalid worksheet schema')
if not isinstance(approval, dict) or set(approval)!= {'digest', 'target', 'lineage', 'serial', 'allowDestruction'}:
raise ValueError('invalid approval schema')
for obj in (report, approval):
for field in ('target', 'lineage'):
if not isinstance(obj[field], str) or not obj[field].strip:
raise ValueError('nonempty identity required')
if type(obj['serial']) is not int or obj['serial'] < 0:
raise ValueError('nonnegative integer serial required')
if type(report['unknownEffects']) is not bool or type(approval['allowDestruction']) is not bool:
raise ValueError('boolean flags required')
if not isinstance(approval['digest'], str) or len(approval['digest'])!= 64 or any(c not in '0123456789abcdef' for c in approval['digest']):
raise ValueError('lowercase SHA256 digest required')
actions = report['actions']
if not isinstance(actions, list) or any(not isinstance(x, str) or x not in {'no-op', 'create', 'update', 'delete', 'replace'} for x in actions):
raise ValueError('unsupported worksheet action')
def assess(report, approval):
validate(report, approval)
checks = {
'digest': digest(report) == approval['digest'],
'target': report['target'] == approval['target'],
'lineage': report['lineage'] == approval['lineage'],
'serial': report['serial'] == approval['serial'],
'known_effects': not report['unknownEffects'],
'destruction_scope': approval['allowDestruction'] or not any(a in {'delete', 'replace'} for a in report['actions']),
}
return {'readyForReview': all(checks.values),
'failedChecks': [key for key, ok in checks.items if not ok],
'executionAuthorized': False}
def fixture(name, report, approval):
return {'id': name, **assess(report, approval)}
def main:
base = {'target': 'preview-west', 'lineage': 'classroom-estate-a', 'serial': 41,
'unknownEffects': False, 'actions': ['update']}
approval = {key: base[key] for key in ('target', 'lineage', 'serial')}
approval.update(digest=digest(base), allowDestruction=False)
fixtures = [fixture('matching-evidence', base, approval)]
for key, value in [('digest', '0' * 64), ('target', 'prod-west'),
('lineage', 'classroom-estate-b'), ('serial', 44)]:
changed = {**approval, key: value}
fixtures.append(fixture('different-' + key, base, changed))
unknown = {**base, 'unknownEffects': True}
fixtures.append(fixture('unknown-effects', unknown, {**approval, 'digest': digest(unknown)}))
destructive = {**base, 'actions': ['replace']}
for allowed in (False, True):
fixtures.append(fixture('replacement-' + str(allowed).lower, destructive,
{**approval, 'digest': digest(destructive), 'allowDestruction': allowed}))
combinations = 0
for bits in product((False, True), repeat=6):
r = deepcopy(base); a = deepcopy(approval)
r['unknownEffects'] = not bits[4]
r['actions'] = ['update'] if bits[5] else ['replace']
a['digest'] = digest(r) if bits[0] else '0' * 64
if not bits[1]: a['target'] = 'prod-west'
if not bits[2]: a['lineage'] = 'classroom-estate-b'
if not bits[3]: a['serial'] = 44
result = assess(r, a)
assert result['readyForReview'] == all(bits)
assert len(result['failedChecks']) == bits.count(False)
assert result['executionAuthorized'] is False
combinations += 1
invalid = [({}, approval), ({**base, 'serial': True}, approval),
({**base, 'unknownEffects': 'false'}, approval),
({**base, 'actions': ['unsupported']}, approval),
({**base, 'actions': 'update'}, approval),
({**base, 'lineage': ''}, approval),
(base, {**approval, 'digest': 'A' * 64}),
(base, {**approval, 'allowDestruction': 1}),
(base, {**approval, 'serial': -1})]
for r, a in invalid:
try: assess(r, a)
except ValueError: pass
else: raise AssertionError('invalid input accepted')
before = deepcopy((base, approval)); assess(base, approval)
assert before == (base, approval)
assert [f['id'] for f in fixtures if f['readyForReview']] == ['matching-evidence', 'replacement-true']
print(json.dumps({'fixtures': fixtures, 'gateCombinations': combinations,
'invalidInputs': len(invalid), 'inputPreserved': True,
'network': False, 'persistentWrites': False, 'terraformExecuted': False,
'authorizationProof': False, 'independentVerification': False,
'scriptSha256': sha256(Path(__file__).read_bytes).hexdigest}, indent=2))
if __name__ == '__main__':
main
The ticket approves digest A for preview-west, but the runner receives B for prod-west. Hold, confirm state identity and review the new plan. Copying B into the ticket without review does not resolve mismatch.
Common pitfalls
Confusing approved commit with executed plan; forcing state without reconciliation; deleting mismatched checksums; importing one object into two managers; using direct deletion to bypass PDB; assuming a freeze eliminates maintenance.
Related topics: Configuration and drift management · Resilience during changes · Governance and RUN handover
Plan identity, state freshness and operational capacity are distinct evidence. Approval must bind to what will happen on the correct target.
Reference: Terraform apply command · Current linked guide; edition date unconfirmed (2026-09-30 inspection)