← Professional Cloud DevOps Engineer: delivery and reliability
10 / 15 · 120 MIN

IaC state and environment lifecycle

Coordinate changes, separate state and permissions, verify defaults and retire temporary environments using explicit evidence.

Identify the state controlled by each run

Before investigating an apply, identify the run and selected state. Record repository, revision, backend, prefix, workspace and executing identity. With the GCS backend, state is an object whose path depends on the prefix and workspace within the selected bucket. Two jobs can use different variables yet address the same state. In a fictional rehearsal, development and production received separate pipelines but retained identical backend configuration. A contested lock was the first symptom; the design problem was missing intentional deployment separation. Confirm whether sharing is deliberate before changing parameters. For independent deployments, agree state organization and permissions with their owners. Maintain a state-recovery approach too; backend documentation recommends object versioning. A different branch name does not establish isolation. Evidence should let another operator reconstruct exactly which resource set the run intended to manage.

Treat locks as coordination information

A lock signals that state access needs coordination. If another job is active, do not use force-unlock merely because the change window is ending. Identify the lock holder and correlate it with run status, logs and its owner. A few seconds without new log lines do not establish abandonment either. When a run terminates abnormally, follow recovery procedure and confirm the writer is no longer acting before handling an orphaned lock. Unlocking does not directly change resources, but it can enable subsequent concurrent changes. Disabling locking or copying state elsewhere also fails to coordinate access to the same remote resources. During an incident, communicate the delay, ongoing operation and condition required for resumption. This helps the window manager more than repeatedly launching applies until one succeeds. Retain the run identification supporting the decision to wait, cancel or recover.

Reconcile state, configuration and emergency changes

Distinguish intended configuration, stored state and objects that actually exist at the provider. Emergency intervention can change only the third. From Terraform 0.15.4, a refresh-only plan lets you review record and output updates reflecting remote changes. Applying that plan neither restores a resource to its old configuration nor automatically edits versioned code. If a resource increased from four units to ten during an incident, document whether ten should remain or four should be restored before resuming normal workflow. A later normal plan can again propose what configuration requests. State separation needs attention too: CLI workspaces use the same backend and do not themselves provide an appropriate credential boundary between teams. For segregated-access requirements, review backend organization and permissions. The name prod identifies context but is not an authorization control. Keep the reconciliation decision visible in the handover record.

Manage temporary environments with a checkable inventory

A deadline in inventory or a label is useful only when a process interprets it. Define who can create, extend and retire an environment, how active work is identified and which data must be retained. This lesson’s local exercise has an explicit fictional rule: propose review only when environment is preview, expiry has been reached, an owner exists, hold and active_lease are false, and the snapshot is zero to five minutes old. These conditions are not an automatic Google Cloud policy. The program deletes nothing, and a candidate still needs the approved operational procedure. A timestamp without a timezone is insufficient for deadline evaluation; the text "false" is not treated as boolean false either. Future or stale snapshots are held for investigation. When reporting to the financial owner, distinguish expired resources, validated candidates and completed deletion. This avoids promising savings from a list still containing environments used by tests or incidents. These inventory fields are not a proposed set of Cloud labels.

Check fleet defaults and effective configuration

A default establishes baseline configuration for eligible new clusters, but does not demonstrate that existing members were updated. Inventory each cluster’s effective version, plan required synchronization and check the results. Feature presence alone also does not prove desired configuration reached the application. For Config Sync, current documentation supports configuring source-of-truth connections in defaults through CLI or Terraform, while the console does not offer that connection as a default. If installation used the console and no package was subsequently configured, do not expect the component to discover the application Git repository. Define source, revision and read authorization, then observe synchronization. For private repositories, an engineer’s access does not replace access by the component’s identity. During RUN handover, record feature status, the revision actually applied and errors preventing convergence. These separate observations make rollout progress and remaining ownership clear.

Prepare tooling without depending on the previous session

Cloud Shell and Cloud Workstations provide development tooling, but their lifecycles need to be understood. In normal Cloud Shell mode, the home directory persists between sessions while the VM is temporary. A manual installation under /usr/local can disappear when the VM changes. Use reproducible preparation and check versions before executing procedures; file persistence is neither a backup nor unlimited retention. In Cloud Workstations, a CPU-bound background process can remain busy without producing events that reset idle timeout. Documentation distinguishes incoming network requests, Start calls and IDE interaction. If long-running work is required, choose an approved path considering timeouts, recovery and retained results. Do not increase memory or privileges without evidence they address the cause. A runbook should allow an operator to start a fresh session and reach the same toolset and authorized context.

Control dependency resolution at the client

CI/CD architecture includes the path used to obtain dependencies. A virtual repository can prioritize its aggregated upstreams, but it does not control requests a client sends directly to another index. In a fictional example, pip queries the internal endpoint and PyPI as an additional index. Private priority at the internal endpoint does not govern the client’s entire selection. Configure the client to use the approved endpoint and pin the intended version; also check which upstreams can supply that version. Larger numeric values mean higher priority, and a tie does not identify a unique origin. These measures reduce a selection gap but do not demonstrate package safety. Dependency approval still requires appropriate origin, review and analysis evidence. To investigate differences between builds, retain effective client configuration, resolved version and observed source alongside application code.

Run the model and prepare the operational decision

The code below uses only Python’s standard library and embedded fictional data. Run it in a trial directory with python3 run.py and interpret results before changing any rule. Alpha, lambda and mu are candidates under the supplied conditions: lambda demonstrates timezone equivalence, and mu sits exactly at the five-minute freshness boundary. Other records demonstrate missing ownership, production scope, future expiry, active lease, hold, stale snapshot, timezone-free date, text boolean and future snapshot. The program checks 64 combinations of six conditions, six time boundaries and five invalid date inputs. It neither queries cloud inventory nor establishes that metadata reflects reality. Add to your reasoning who collects and validates that evidence. Then resolve the expired-environment case with an active apply. Explain to the window manager why expiry does not establish an abandoned lock and what evidence is missing before cleanup can be decided.

"""Original fictional inventory review. No provider calls or deletion operations."""
from datetime import datetime, timezone, timedelta
from itertools import product
from pathlib import Path
import hashlib
import json
import platform

NOW = datetime(2026, 10, 6, 12, tzinfo=timezone.utc)


def instant(value):
 if not isinstance(value, str):
 return None
 try:
 parsed = datetime.fromisoformat(value)
 except ValueError:
 return None
 if parsed.tzinfo is None or parsed.utcoffset is None:
 return None
 return parsed.astimezone(timezone.utc)


def classify(row, now=NOW):
 """Select a review candidate only under the six explicit exercise rules."""
 expires = instant(row.get('expires_at'))
 observed = instant(row.get('observed_at'))
 checks = {
 'preview_only': row.get('environment') == 'preview',
 'expired': expires is not None and expires <= now,
 'owner_known': isinstance(row.get('owner'), str) and bool(row['owner'].strip),
 'no_hold': row.get('hold') is False,
 'no_active_lease': row.get('active_lease') is False,
 'fresh_snapshot': observed is not None and timedelta(0) <= now - observed <= timedelta(minutes=5),
 }
 reasons = [key for key, passed in checks.items if not passed]
 return {'id': row['id'], 'review_candidate': not reasons, 'reasons': reasons}


def main:
 base = dict(id='alpha', environment='preview', expires_at='2026-10-06T11:00:00Z',
 owner='team-recon', hold=False, active_lease=False,
 observed_at='2026-10-06T11:58:00Z')
 fixtures = [base,
 dict(base, id='beta', owner=''),
 dict(base, id='gamma', environment='production'),
 dict(base, id='delta', expires_at='2026-10-06T13:00:00Z'),
 dict(base, id='epsilon', active_lease=True),
 dict(base, id='zeta', hold=True),
 dict(base, id='eta', observed_at='2026-10-06T11:54:59Z'),
 dict(base, id='theta', expires_at='2026-10-06T11:00:00'),
 dict(base, id='iota', hold='false'),
 dict(base, id='kappa', observed_at='2026-10-06T12:00:01Z'),
 dict(base, id='lambda', expires_at='2026-10-06T13:00:00+01:00'),
 dict(base, id='mu', observed_at='2026-10-06T11:55:00Z')]
 results = [classify(row) for row in fixtures]
 assert [r['id'] for r in results if r['review_candidate']] == ['alpha', 'lambda', 'mu']
 named_checks = ['fixture_candidates', 'production_excluded', 'owner_required',
 'active_lease_excluded', 'hold_excluded', 'stale_snapshot_excluded',
 'naive_timestamp_rejected', 'string_boolean_rejected',
 'future_snapshot_rejected', 'timezone_equivalence', 'inclusive_expiry',
 'inclusive_freshness']
 assert results[2]['reasons'] == ['preview_only']
 assert results[1]['reasons'] == ['owner_known']
 assert results[4]['reasons'] == ['no_active_lease']
 assert results[5]['reasons'] == ['no_hold']
 assert results[6]['reasons'] == ['fresh_snapshot']
 assert results[7]['reasons'] == ['expired']
 assert results[8]['reasons'] == ['no_hold']
 assert results[9]['reasons'] == ['fresh_snapshot']
 assert instant('2026-10-06T13:00:00+01:00') == NOW
 assert results[10]['review_candidate']
 assert results[11]['review_candidate']
 truth_cases = 0
 for flags in product([False, True], repeat=6):
 row = dict(base, environment='preview' if flags[0] else 'production',
 expires_at='2026-10-06T11:00:00Z' if flags[1] else '2026-10-06T13:00:00Z',
 owner='team' if flags[2] else '', hold=not flags[3],
 active_lease=not flags[4],
 observed_at='2026-10-06T11:59:00Z' if flags[5] else '2026-10-06T11:54:00Z')
 result = classify(row)
 assert result['review_candidate'] == all(flags)
 assert len(result['reasons']) == flags.count(False)
 truth_cases += 1
 expiry_cases = 0
 for seconds in [-3600, -1, 0, 1, 60, 3600]:
 row = dict(base, expires_at=(NOW + timedelta(seconds=seconds)).isoformat)
 assert classify(row)['review_candidate'] == (seconds <= 0)
 expiry_cases += 1
 for invalid in [None, 17, '', 'not-a-date', '2026-99-99T12:00:00Z']:
 assert instant(invalid) is None
 print(json.dumps({'scriptSha256': hashlib.sha256(Path(__file__).read_bytes).hexdigest,
 'runtime': platform.python_version, 'checks': len(named_checks),
 'checkNames': named_checks, 'fixtures': results, 'truthCases': truth_cases,
 'expiryCases': expiry_cases, 'invalidTimestamps': 5, 'network': False,
 'persistentWrites': False, 'deletions': 0, 'vendorExecution': False,
 'independentVerification': False}, indent=2))


if __name__ == '__main__':
 main
IN PRACTICE

Two jobs address the same state while a cleanup routine considers an environment whose deadline expired but which still has active work.

Common pitfalls

Forcing active locks, confusing refresh-only with resource restoration, using names as access control and assuming component installation configures its source.

Related topics: State and drift management · Development platforms and GitOps · Obsolescence, decommissioning and FinOps

Take this idea with you

Before changing or retiring an environment, identify state, identity, effective configuration and ongoing work, then demonstrate the intended outcome.

Create account

Reference: GCS backend · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.