← Professional Cloud Architect: architecture and operations
10 / 14 · 120 MIN

Infrastructure: routes, recovery and pipelines

Configure hybrid connectivity, recovery, retention and capacity with evidence for RUN handover.

Prove the path the business uses

An active tunnel establishes only part of the path. Before accepting a migration, write a matrix containing source, destination, prefix, port, protocol and expected result. Distinguish routes Cloud Router learns from routes it advertises to a peer. Custom-only can exclude previously advertised subnets; a session in custom mode can also stop inheriting router configuration. Retain effective advertisements per session and compare them with the matrix. For a funds application, a missing prefix can affect batch processing even with BGP established. Check the return direction too. This exercise is a document analysis: choose one flow, identify where each route must exist and describe the evidence you would request from the network team. A single ping does not replace testing the application protocol or show that every required flow works.

Review regional scope before changing the firewall

A VM in a new region may lack the dynamic route used by existing VMs. Regional mode limits processing of learned routes to the relevant region; global mode allows consideration of paths from other regions of the same VPC. That choice does not mean automatic transit between arbitrary networks. In the change plan, record which regions need the on-premises service, expected paths and each connection’s dependencies. Changing a firewall without a route creates no connectivity. Seeing a prefix on a router is insufficient too: confirm the route applicable to the resource and the return path. For a migration window, request before-and-after evidence and define rollback for unexpected paths. Explain the intended reach, the observation that demonstrates it and who confirms effects on shared applications.

Prevent recovery from amplifying failure

A load-balancing health check determines where traffic goes; autohealing can recreate a VM. Use criteria suited to each consequence. Blocked probes can classify a healthy application as failed. A check overly sensitive to brief pauses can make recovery cause a longer interruption. Initial delay helps during startup, but success can end it early: do not return healthy merely because the process started if an essential phase remains. In the MIG scenario, start with probe logs and application state, retain evidence and fix observation before attributing failure to capacity. Design a rehearsal including slow startup, a transient pause and persistent failure. For each, state whether you expect traffic withdrawal, waiting or recreation, and which result would halt rollout. Record recovery duration as well as the final green status.

Inventory data before promising savings

A lifecycle update can take up to 24 hours, during which the previous rule can still act. Do not treat update acknowledgement as immediate protection. Age, holds, retention, noncurrent versions and soft delete are also different inventory dimensions. In a versioned bucket, removing the live version can preserve recoverable bytes and cost. For FINOPS, prepare a table per dataset covering live volume, history, applicable protection, transition rule and expected deletion evidence. The exercise does not ask you to remove retention to meet a financial target; it asks you to explain why the forecast changed. If a team claims everything was deleted because normal listing is empty, request the relevant version inventory. Test policy on rehearsal data before applying it to the intended set, recording which objects should remain protected.

Separate recovered disk from recovered service

A regional disk reduces certain zonal dependencies but does not execute the entire recovery plan. If the original VM cannot detach, an authorized procedure can use force-attach on the recovery VM; after that operation, Compute Engine prevents the original VM from writing to that disk. This guarantee neither undoes external calls nor confirms application startup. The runbook should identify the disk, target VM, mounting, application checks and traffic restoration. Define decision points before accepting work again. During rehearsal, infrastructure supplies volume-access evidence and APS confirms a functional flow using known references. Measure time until usable service. Recording only attach time omits phases that can dominate recovery and create a false impression that the objective was met. Confirm who owns each transition before the window starts.

Provision with explicit dependencies

An uneditable instance template can contain a script downloading latest; object immutability therefore does not guarantee repeatable installed bytes. Pin versions and create a new template for a change. In GKE Standard, inspect requests and scheduling constraints when Pods are Pending, even if measured CPU is low. IP space can also prevent new nodes or Pods: increasing maximum nodes does not fix an exhausted range. In Cloud Run, private-ranges-only cannot meet an all-destinations-through-VPC requirement; all-traffic still needs outbound-path validation. Bring these dependencies into one capacity review: installed version, reserved resources, addressing, quotas and network path. Request a concrete measurement or configuration for each. A generic statement that autoscaling is enabled does not establish that the application can grow within the required window or remain operable after replacement.

Make ML pipelines reproducible and authorized

Current documentation places pipelines under Gemini Enterprise Agent Platform. Use the current source without assuming an old URL defines the product name. Step caching depends on the declared interface, including parameters, artifact IDs and component specification. If current.csv changes without changing that identity, the application can reuse a result no longer matching the intended data. Represent the version or disable the affected step’s cache when needed. If execution fails while writing artifacts, identify the runtime service account: the submitting engineer’s permissions are not automatically inherited. For the exercise, build a record containing dataset version, component, cache decision, execution identity and resource permissions. That record should explain why two outputs differ without relying only on the job’s display name. Retain evidence of what actually ran and what was reused.

Select AI capabilities and close acceptance

Choose an API by the task and required output. In a Cloud Vision prototype, DOCUMENT_TEXT_DETECTION provides dense-text structure; it does not validate a financial instruction. Other solutions can be appropriate, including Document AI for specific document requirements. Likewise, a model’s presence in Model Garden does not replace review of warnings, execution mode and internal criteria. A suspicious model can remain technically deployable. As a final exercise, prepare a go/no-go decision for the hybrid scenario: list both observed causes, proposed changes, required evidence and an alternative conditional on approved rollback. Then add an ML dependency using the batch files and explain which data identity and permissions must appear in handover. This is a document-based analysis and does not claim execution on Google services. Make each acceptance statement traceable to an observation.

IN PRACTICE

A migration has an active tunnel, missing prefix and blocked probes. Acceptance requires fixing routing and observation, then testing service.

Common pitfalls

Confusing active BGP with reachable service, recreation with recovery, non-live objects with deleted data and job submitters with runtime identity.

Related topics: Hybrid migration · Capacity and FINOPS · RUN handover

Take this idea with you

Connect configuration, identity and observation: each guarantee has a scope and each acceptance decision needs evidence within that scope.

Create account

Reference: Cloud Router advertised routes · Current linked standard guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.