← Professional Cloud Architect: architecture and operations
20 / 25 · 120 MIN

Signals, recovery and shift handoff

Interpret telemetry, distinguish recovery milestones and transfer operational decisions with clear evidence and ownership.

Start with the operational question and unit

A dashboard should support a concrete decision: investigate delay, stop an expansion, confirm recovery or track a dependency. Before choosing colors, define population, period and unit. In the CPU exercise, the three samples are 10%, 10% and 90%. A mean of about 36.7% does not contradict the 90% peak; it summarizes another property. Inspecting the maximum locates the high sample but does not prove it lasted throughout the window or caused an incident. When discussing with development and APS, ask for the link between signal, hypothesis and next check. Retain data capable of contradicting the hypothesis instead of choosing only the most favorable aggregation. In a reporting service, the user can still be waiting for a file while infrastructure averages look acceptable. Combine technical health with a representative functional flow rather than treating one indicator as complete evidence of availability.

Distinguish absence, zero and notification suppression

A zero-error measurement is meaningful only within a known population and collection path. If request and error counters stop reporting, replacing both with zero does not create a valid success measurement. In this lesson’s first case, the functional probe keeps failing while the dashboard hides absence. Keep acceptance pending and investigate both service and telemetry. A second trap occurs at startup: an absence policy for a new heartbeat needs to observe an initial measurement; never having alerted does not establish that the producer works. Rehearsing reporting and interruption belongs in handover. Finally, a snooze can suppress notifications during maintenance. That silence is a configuration, not recovery. Record suppression period and scope, maintain observation suitable for the window and confirm behavior before reporting that the service has returned to normal. A closed monitoring alert likewise needs interpretation in the context of the actions that changed its state.

Read a log without inventing duration or cause

The original exercise fragment has timestamp at 09:10 and receiveTimestamp at 09:14, with a delivery_failed event for job report-17. Assuming the source clock is correct, the event is marked four minutes before Logging receives it. Write three statements: what was observed, what remains unknown and what you would check next. A suitable answer distinguishes event time from platform arrival and asks for evidence about collection, buffering and transport. It cannot conclude that the job ran for four minutes because no field represents its start. Nor are there two errors merely because two times appear. To investigate a release window, retain both time axes and the job identifier. Avoid adding a unique request_id as a metric label: that multiplies series. Keep bounded dimensions in metrics and correlatable detail in logs, with access and retention appropriate to the context. Renaming or fully hashing unique values does not solve their cardinality.

Follow the execution chain and work age

When A calls B, two exported spans do not guarantee a connected distributed trace. In the scenario, B ignores incoming context and always creates a new root. Correct context propagation and extraction, then confirm trace ID, span ID and parent span ID along the chain; increasing sampling does not fix this structural failure. For an asynchronous flow, add another question: how long has the oldest work been waiting? A subscription can retain only three unacknowledged messages while one grows older and new arrivals are processed. Low volume does not establish timely completion. Investigate persistent messages, processing failures, acknowledgements and capacity before concluding that every consumer needs more CPU. The exercise excludes filters and transformations to make the observation clear. In a real system, document those settings because they can change backlog interpretation. Do not acknowledge messages merely to improve a chart: rejection, quarantine and reprocessing need an explicit decision about the work.

Compare request-based and windows-based indicators

The guided table contains ten eligible windows with no missing measurements. In the first nine, each window has 100 successes from 100 requests. The last has zero successes from ten requests. A good window requires at least 99% success within that window. Classify each before calculating: nine are good and one is bad, giving a windows-based SLI of 9/10, or 90%. Now add requests: there are 900 successes from 910 requests, about 98.90%. Neither calculation corrects the other; they measure different units. Explain which matches the agreed objective and avoid choosing the more favorable measure afterwards. As an extension, change the last window to 99 successes from 100 requests: it passes exactly at the threshold without favorable rounding. If you remove a window’s data, you can no longer use the exercise rule without defining missing-data handling and eligibility.

Inventory what each recovery mechanism protects

Complete recovery depends on more than a successful backup status. Backup for GKE protects manifests and supported volumes within configured scope; an image reference does not store its layers, and a Cloud SQL connection does not include the external database. In rehearsal, ask where the recoverable image resides, how the database is protected and how dependencies will be validated in sequence. In Cloud Storage, restoring a soft-deleted object creates a new live version. A consumer pinned to the earlier generation needs reference reconciliation and expected-content testing. In Cloud Run, retaining minimum instances does not guarantee that the same process never restarts; a checkpoint required for resumption should not exist only in memory. For each dependency, record the protected object, recoverable point, access identity, restore action and functional criterion. This makes resilience discussions verifiable and exposes gaps before a real incident, rather than after a successful component restore is mistaken for service recovery.

Measure recovery through the agreed functional outcome

In the second case, the new funds-recovery instance exists at 14:12 and passes data validation at 14:18. At 14:25, the client still points to funds-live and functional testing fails. The exercise defines RTO completion as restoring functional service, so the clock cannot stop at 14:18. That instant remains a useful milestone with a precise name: database validated. The next plan must coordinate write destinations, access, connection configuration, reconciliation and application testing. PITR does not automatically rewrite clients; deleting the source is not a mechanism for selecting the new instance either. Retain recovery options while transition remains unvalidated and record authorized decisions. At the committee, present time spent on each stage and what remains before closing the objective. Distinguishing a ready component from recovered service helps improve the next rehearsal based on the actual observed path rather than an assumed instantaneous client transition.

Hand responsibility and evidence to the next shift

A long incident can cross teams, languages and time zones. Prepare a living record of known impact, open hypotheses, performed changes, decisions and next actions. When handing over from Lisbon to another team, review state with the incoming lead, obtain explicit acceptance and communicate who coordinates from that moment. A message sent without acknowledgment does not establish transfer of command. For practice, ask a colleague to explain the latest decision and the condition for changing it using only the handoff record. If they cannot distinguish a validated database from a recovered service, or missing metrics from zero errors, improve the document and repeat. This lesson’s summary is to preserve meaning: what a signal measures, what a mechanism protects and who owns the next decision. Connect this work with change governance, data reconciliation and release acceptance. These original examples are exercises; they do not constitute execution of recovery in Google Cloud.

{"timestamp":"2026-10-06T09:10:00Z","receiveTimestamp":"2026-10-06T09:14:00Z","severity":"ERROR","jsonPayload":{"job":"report-17","event":"delivery_failed"}}
IN PRACTICE

The recovered database passes validation at 14:18, but the application still points to the source and fails at 14:25. Service recovery remains pending.

Common pitfalls

Replacing absence with zero, inferring duration from log receipt, hiding peaks in a mean, ignoring images and external databases, and stopping RTO before functional acceptance.

Related topics: Release acceptance and observability · Data reconciliation and recovery · Incident governance and coordination

Take this idea with you

Define the indicator unit, confirm recovery scope and hand the next shift a decision that can be explained and continued.

Create account

Reference: Filtering and aggregation: manipulating time series · Current linked standard guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.