← Professional Cloud Security Engineer: controls and evidence
11 / 23 · 135 MIN

Data protection, keys and AI context

Connect authorized reading, recovery dependencies and AI boundaries in production decisions.

Start with the read that should be permitted

In a data-access incident, start by writing who is asking, which operation they need and which data it concerns. “The team needs report access” hides different decisions: querying portfolio operations, reading full identifiers, exporting results and sharing them with another group. In a fictional funds operations service, an analyst might need to investigate reconciliation differences without receiving every customer's identifying data. This exercise does not describe internal BNP Paribas procedures. In BigQuery, reason about rows and columns separately. Row restriction selects visible operations; column protection limits values within those operations. In the policy-tag workflow, reading a protected column also requires tag authorization alongside dataset access. Before fixing a denial, establish approved need and the effective query identity. A test with the administrator account can succeed while still failing to explain the analyst's problem. Include the intended read scope in the incident record so the remediation has a clear boundary. When reviewing row policies, consider every applicable grant. A retained broad policy can continue permitting rows despite a new narrower policy. Prepare two fictional operation sets, one allowed and another that should be rejected. Record expected results before testing. Review is incomplete if it only demonstrates that a query returns data; it needs to check that the user receives the right set and that old exceptions have been addressed. This makes successful access and preserved separation equally visible during acceptance.

Migrate permissions without expanding disclosure

Consider a legacy bucket where each vendor receives its files through object ACLs. Moving to uniform bucket-level access requires mapping those relationships to the intended IAM design. Giving every former reader whole-bucket read access can keep the batch working while breaking separation between vendors. A technical success indicator alone does not capture that effect. First draw the matrix of vendor, object set and authorized operations, including who creates new files and who supports failed deliveries. Separate sets with different requirements through a supported design, for example appropriate bucket groupings. Before the change window, confirm consumers, identities, policies and processes for publishing new files. An object created after migration also needs the correct scope. Include negative tests: the identity used by vendor A must still be unable to read a file intended only for B. Use fictional names and content in rehearsal evidence, with expected outcomes agreed before execution. Public access prevention addresses a different scope and does not block Signed URLs. A conclusion such as “the bucket is private, so nobody can share access” therefore remains unproven. Also review who issues those URLs, the object covered and the delivery process. At go-live review, present unresolved dependencies, test results and decision owners. If separation is unproven, the plan may need correction or deferral; expanding permissions should not become the automatic contingency mechanism.

Interpret findings and preserve context

An inspection report describes one execution, not the entire reality of the data. Retain source, time range, detector configuration, exclusions and inspected scope. If 100 of 50000 records were examined, no findings does not establish that the remaining records contain no sensitive information. Apply the same care to uninspected fields and information types that configuration did not seek. A sample can guide investigation without supporting a definitive classification of the entire set or future additions to it. In the fictional reconciliation process, compare usefulness requirements with exposure reduction. A transformed identifier can enable necessary joins between two tables. If transformation uses a different context tweak in each pipeline, the same value and key can produce different tokens. Before concluding that corruption occurred, compare configuration, normalization and the actual supplied context. Document which correlations are intended and which should remain separate, so a technically successful join does not conceal a changed data-use boundary. The correction is not to make every transformation identical without examining the requirement. Different context can intentionally limit correlation between portfolios. If the join requirement changed, the decision should include data owners and downstream consumers. Rehearsing with synthetic values allows matching checks without placing real identifiers in the ticket. In the summary, distinguish detection, classification and transformation: each answers a different question, and none removes the need to know who can use the result.

Treat secret rotation as a change

A rotation schedule is useful only when the complete process has an owner. Secret Manager sends the SECRET_ROTATE notification through Pub/Sub; a subscriber and its workflow need to perform the necessary actions. For an external integration, identify who changes the credential in its source system, publishes the new version, updates consumers and confirms operation. Receiving the message is evidence of one stage, not proof that every stage finished successfully or that no old consumer remains. Design the sequence before the window. If the external system temporarily accepts two credentials, a transition strategy may differ from one for a system accepting only one. That capability is a scenario condition requiring confirmation in the actual system. Define what happens if creation succeeds but publication fails, or if some instances still use the old version. Rolling back code might not restore a credential already invalidated by the provider. Record those dependencies in the recovery plan rather than relying on a generic rollback statement. Numeric version references make release dependencies explicit and allow rehearsal before adoption. Avoid consuming an incompatible version without a decision merely because latest changed. Finally, examine copies: a supported integration might inject the secret through an environment variable, but a library logging the environment creates exposure in logs. Production handover should include diagnostics without secret values, failure signals and an owner for consumers that did not complete transition.

Separate key availability, integrity and identity

When a cryptographic operation fails, distinguish version state, permissions, selected key and data integrity. A version restored during its scheduled destruction period returns to DISABLED. Re-enabling must be evaluated before using it again. Do not confuse successful administrative restoration with confirmation that an archive can already be recovered. Relate the observed state to the intended operation and the responsible team's authorization, and record which next action has actually been approved. In a KMS operation response, checking bytes and checking the resource are different steps. If the ciphertext checksum mismatches, discard that result and handle retries within limits. If bytes are consistent but name identifies an unexpected key, investigate resource identity. Replacing metadata with the desired name does not change the key that produced the result. Retain diagnostic evidence without plaintext or secret material. A successful transport response should not cause either independent check to disappear from the recovery procedure. For a CMEK-supported Agent Platform resource, the administrator initiating creation and the service agent using the key are different identities. If failure belongs to the service agent, giving the person more privileges does not fix that dependency. Confirm the documented identity, key and location for the particular resource. For the technical PM, useful handover is a matrix of operation, identity, resource, evidence and owner, rather than just a screenshot of an administrative role grant.

Plan recovery beyond ciphertext

In an envelope-encryption design, a DEK can encrypt an archive and then be protected by a KEK. Recovery needs the archive, protected DEK and relevant references and permissions. Keeping only the KEK name does not reconstruct a discarded DEK. Draw the recovery package and ask which components would survive failure of the originating service. This conceptual exercise implements no cryptography and does not recommend inventing algorithms. Its purpose is to expose dependencies that a simple file-backup count would miss. AAD adds a context dependency. Decryption requires the same value used for encryption. In the fictional case, the team included a contract version in AAD and later replaced that metadata during archive migration. Key availability does not resolve lost original context. Test whether the process preserves or reproduces that value without confusing it with the current contract version. Also record the chosen location: a global key does not establish confinement to one specific region, even if a consumer currently runs there. For data in use, attestation supplies evidence for a verifier to compare against policy. A valid signature over an old observation does not establish freshness. If policy requires a fresh challenge, check that binding, expected measurements and the decision of the key-release service. The PM should request sufficient evidence for the approved criterion, including expected failures. Do not extrapolate one positive test to every Confidential VM technology or to overall application security.

Control what enters AI context

A search assistant introduces two identities that must remain distinct: the technical identity reading the source and the person receiving the answer. If the service account reads every portfolio, that does not establish that every user may receive every document. Authorize content for the request before placing it in model context. Filtering only citations in the answer happens too late to prevent content being transmitted for processing. Include this boundary in the data-flow review and acceptance criteria. The calling path also matters. An agent built to trust a frontend can depend on that frontend to manage users and sessions. Giving a new client direct endpoint access requires reviewing controls in the receiving code. Do not treat a client-supplied user identifier as sufficient authorization evidence. Write one permitted case and another where the client attempts to use somebody else's session, without using real customer data in rehearsal. Endpoint access should be assessed alongside the application's own trust assumptions. Retrieved text can contain instructions attempting to divert model behavior. Document relevance does not give it authority to export a report or call a tool. Action authorization still requires a server-side control. Model Armor can help inspect content, but integration must define outcomes and failures. Here, fictional policy requires a valid permitting inspection response; a timeout leaves that condition unmet. This is an explicit example rule, not a claim about the default behavior of every integration. Make that distinction clear when adapting the case to a real service.

Exercise: select documents with bounded evidence

The Python exercise runs locally and uses only fictional metadata. It reads no documents, calls no models and does not evaluate Google IAM. Its teaching policy requires the same tenant, a group match, and complete recent evidence on both sides. Groups are scoped to the tenant; the same group name in another tenant grants no access. The function first classifies eligibility and only then ranks eligible documents by score, using the identifier to break ties. These rules are deliberately explicit application assumptions. Before running python3 run.py, predict the outcome: one permitted document scores 0.70 and another from a different tenant scores 0.99. Only the first should be selected. Then change ACL age: the limit is inclusive, so age 10 passes with max_age=10, but age 11 does not establish freshness. Incomplete, undated or future-dated evidence leaves the document unresolved. A known tenant mismatch remains recorded even when ACL evidence is also missing. Check both reasons instead of hiding a known failure behind uncertainty elsewhere. Read selectedIds as this model's result, not proof of real access. Diagnostic identifiers and reasons are for the exercise operator, not an end-user response contract. Complete the activity by writing a note with the excluded document, known condition and next required collection. The summary is straightforward: authorized data, preserved cryptographic context and bounded AI actions require different decisions. Evidence should let the team distinguish them during incidents and before production handover.

"""Original offline teaching model. No cloud IAM or real retrieval evaluation.

Policy: same tenant, matching group, complete and fresh ACL/membership evidence.
Freshness ages 0..max_age inclusive. Unknown or stale evidence withholds eligibility.
Diagnostic IDs are for the exercise operator, never an end-user response contract.
"""
from copy import deepcopy
from hashlib import sha256
from itertools import permutations, product
from math import isfinite
from pathlib import Path
import json


def number(value, label):
 if type(value) not in (int, float) or not isfinite(value):
 raise ValueError(label + ' must be a finite number')
 return value


def string(value, label):
 if not isinstance(value, str) or not value.strip:
 raise ValueError(label + ' must be a nonempty string')
 return value


def groups(value):
 if not isinstance(value, list):
 raise ValueError('groups must be a list')
 for group in value:
 string(group, 'group')
 if len(value)!= len(set(value)):
 raise ValueError('duplicate group')
 return set(value)


def evidence(record, now, max_age):
 if type(record.get('complete')) is not bool:
 raise ValueError('complete must be boolean')
 at = record.get('checkedAt')
 if at is not None:
 number(at, 'checkedAt')
 unknown = []
 if not record['complete']:
 unknown.append('incomplete')
 if at is None:
 unknown.append('missing-time')
 elif at > now:
 unknown.append('future-time')
 elif now - at > max_age:
 unknown.append('stale')
 return unknown


def select_documents(caller, documents, now=100, max_age=10, top_k=2):
 number(now, 'now')
 number(max_age, 'max_age')
 if max_age < 0 or type(top_k) is not int or top_k < 0:
 raise ValueError('invalid age or selection limit')
 if not isinstance(caller, dict) or not isinstance(documents, list):
 raise ValueError('caller and documents shape')
 tenant = string(caller.get('tenant'), 'caller tenant')
 caller_groups = groups(caller.get('groups'))
 caller_unknown = evidence(caller, now, max_age)
 seen, decisions, eligible = set, [], []
 for doc in documents:
 if not isinstance(doc, dict):
 raise ValueError('document must be an object')
 identifier = string(doc.get('id'), 'document ID')
 if identifier in seen:
 raise ValueError('duplicate document ID')
 seen.add(identifier)
 doc_tenant = string(doc.get('tenant'), 'document tenant')
 doc_groups = groups(doc.get('groups'))
 score = number(doc.get('score'), 'score')
 acl_unknown = evidence(doc, now, max_age)
 unknown = ['membership:' + x for x in caller_unknown]
 unknown += ['acl:' + x for x in acl_unknown]
 failures = []
 if tenant!= doc_tenant:
 failures.append('tenant-mismatch')
 # Incomplete or stale group lists cannot prove either current grant or denial.
 if not caller_unknown and not acl_unknown and not caller_groups & doc_groups:
 failures.append('no-group-grant')
 status = 'rejected' if failures else ('unresolved' if unknown else 'eligible')
 decisions.append(dict(id=identifier, status=status, knownFailures=failures,
 unresolvedEvidence=unknown))
 if status == 'eligible':
 eligible.append((score, identifier))
 ranked = [identifier for _, identifier in sorted(eligible, key=lambda x: (-x[0], x[1]))]
 return dict(selectedIds=ranked[:top_k], eligibleIds=ranked,
 decisions=sorted(decisions, key=lambda x: x['id']),
 realAccessProven=False, productionAuthorized=False)


def main:
 caller = dict(tenant='A', groups=['ops'], complete=True, checkedAt=100)
 doc = dict(id='a', tenant='A', groups=['ops'], complete=True, checkedAt=100, score=.7)
 fixtures = []

 def case(name, docs, expected, actor=None, **kwargs):
 actor = caller if actor is None else actor
 original = deepcopy((actor, docs))
 result = select_documents(actor, docs, **kwargs)
 assert result['selectedIds'] == expected, name
 assert (actor, docs) == original
 fixtures.append(dict(id=name, **result))
 return result

 case('authorized-before-ranking', [doc, dict(doc, id='b', tenant='B', score=.99)], ['a'])
 r = case('unknown-acl', [dict(doc, complete=False)], [])
 assert r['decisions'][0]['status'] == 'unresolved'
 case('stale-acl', [dict(doc, checkedAt=89)], [])
 case('fresh-boundary', [dict(doc, checkedAt=90)], ['a'])
 case('unknown-membership', [doc], [], dict(caller, complete=False))
 r = case('empty-membership', [doc], [], dict(caller, groups=[]))
 assert r['decisions'][0]['knownFailures'] == ['no-group-grant']
 case('future-evidence', [dict(doc, checkedAt=101)], [])
 case('tie-order', [dict(doc, id='b'), doc], ['a'], top_k=1)
 r = case('known-failure-and-unknown', [dict(doc, tenant='B', complete=False)], [])
 assert r['decisions'][0]['knownFailures'] == ['tenant-mismatch']
 assert r['decisions'][0]['unresolvedEvidence'] == ['acl:incomplete']
 case('empty-inventory', [], [])
 case('zero-selection-limit', [doc], [], top_k=0)
 count = 0
 for same_tenant, member, complete_caller, complete_acl, timestamp in product(
 (False, True), (False, True), (False, True), (False, True), (None, 89, 90, 100, 101)):
 actor = dict(caller, complete=complete_caller, groups=['ops'] if member else [])
 item = dict(doc, tenant='A' if same_tenant else 'B', complete=complete_acl, checkedAt=timestamp)
 result = select_documents(actor, [item])
 expected = same_tenant and member and complete_caller and complete_acl and timestamp in (90, 100)
 assert result['selectedIds'] == (['a'] if expected else [])
 decision = result['decisions'][0]
 assert ('tenant-mismatch' in decision['knownFailures']) == (not same_tenant)
 if not same_tenant:
 assert decision['status'] == 'rejected'
 count += 1
 docs = [doc, dict(doc, id='b', score=.8), dict(doc, id='c', tenant='B', score=.99)]
 expected = select_documents(caller, docs)
 permutation_count = 0
 for permutation in permutations(docs):
 assert select_documents(caller, list(permutation)) == expected
 permutation_count += 1
 invalid = [
 (dict(caller, complete='true'), [doc], {}),
 (dict(caller, groups='ops'), [doc], {}),
 (dict(caller, groups=['ops', 'ops']), [doc], {}),
 (dict(caller, tenant=''), [doc], {}),
 (dict(caller, checkedAt=True), [doc], {}),
 (caller, [dict(doc, score=float('nan'))], {}),
 (caller, [dict(doc, score=float('inf'))], {}),
 (caller, [dict(doc, score=True)], {}),
 (caller, [doc, doc], {}),
 (caller, [dict(doc, id='')], {}),
 (caller, [dict(doc, groups=[5])], {}),
 (caller, [dict(doc, checkedAt='100')], {}),
 (caller, [dict(doc, complete=None)], {}),
 (caller, [doc], dict(top_k=True)),
 (caller, [doc], dict(top_k=-1)),
 (caller, [doc], dict(max_age=-1)),
 (caller, [doc], dict(now=float('inf'))),
 (caller, [None], {}),
 ]
 for actor, items, args in invalid:
 try:
 select_documents(actor, items, **args)
 except ValueError:
 pass
 else:
 raise AssertionError('invalid input accepted')
 print(json.dumps(dict(labId='pcse-retrieval-scope', fixtures=fixtures,
 stateCombinations=count, inputPermutations=permutation_count,
 invalidInputs=len(invalid), inputPreserved=True, orderIndependent=True,
 network=False, cloudExecuted=False, persistentWrites=False,
 scriptSha256=sha256(Path(__file__).read_bytes).hexdigest), indent=2))


if __name__ == '__main__':
 main
IN PRACTICE

A migration preserves batch operation but expands reads; restoration leaves a key disabled; an assistant finds documents outside its tenant. Each result needs its own decision.

Common pitfalls

Confusing a sample without findings with complete classification, a notification with executed rotation and document relevance with authorization.

Related topics: Workload identities and authorization · Data continuity and recovery · AI application security

Take this idea with you

Define who reads, preserve the context needed for recovery and authorize data and actions before delivering them to the model.

Create account

Reference: Introduction to column-level access control · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.