← Professional Data Engineer: pipelines and data decisions
14 / 18 · 135 MIN

Change contracts, security and consumers

Assess consumer impact, check security and prepare migrations with acceptance and fallback criteria.

1. Define the change and acceptance boundary

An architecture change can preserve table names while breaking a business process. In a fictional APS example, a team replaces integer amounts in cents with decimal amounts in euros and adds a risk classification. The daily dashboard still opens, but financial close interprets the old unit. Before scheduling the window, identify producers, consumers, units, grain, identifiers and acceptance owners. These examples do not represent BNP Paribas procedures. Write the existing and proposed contracts separately. Record each consumer's requirements, including whether additional fields are accepted. A monthly exporter may be absent from one week's tests; inactivity does not prove that a dependency is absent. Combine a technical inventory, observation and confirmation by owners. For each gap, name the person who must resolve it and explain its effect on the decision. A usable delivery includes a verifiable outcome: close receives the expected movements, calculates values in the correct unit and meets the agreed deadline. A well-formed file provides only part of that evidence. Separate contract acceptance, data reconciliation and application behavior. This distinction helps locate failures without declaring success merely because an execution finished.

2. Change security without losing scope

An external application with a compatible identity provider can use Workload Identity Federation to avoid distributing a long-lived service account key. Define the accepted external identity, required attributes and resource permissions. Federation addresses how credentials are presented; it does not replace decisions about least privilege. Also test a similar identity that should be rejected. VPC Service Controls and IAM perform different checks. A call allowed by the perimeter still needs IAM authorization. In dry-run mode, perimeter violations are recorded without enforcing that configuration. Use those records to identify legitimate paths and adjust specific rules before enabling enforcement. Do not interpret an absence of blocks as validation of prohibited access. Keep positive and negative tests with reproducible context. For a regional BigQuery migration, check the location rules for the CMEK key and taxonomy. A regional key must match the applicable regional dataset location; membership of the same project is insufficient. The policy-tag taxonomy must exist in the destination table's location. Matching names do not prove matching identities or permissions. Keep security decisions with the migration plan so that RUN can maintain the outcome.

3. Distinguish an accepted schema from correct data

The service and the consumer can apply different rules. In an existing BigQuery table, a new top-level REQUIRED column cannot be added directly as though it were NULLABLE. Plan a supported evolution, value validation and consumer transition. Do not generalize this restriction to every nested field or every table reconstruction method; consult the specific operation. Making a field nullable may be supported by storage and still break a reader that requires a value. Adding a column may suit a reader that selects fields by name while breaking one that rejects unknown fields. Test both perspectives. Use examples where the producer is valid but the consumer fails, to avoid confusing service validation with integration compatibility. The same structure can also conceal a change of grain: one row per movement becomes one row per customer per day. This lesson's exercise does not detect that change. Add reconciliation of counts, keys, totals and business cases. When changing sensitive-data detection thresholds, measure false negatives and false positives; raising minimum likelihood can miss genuine occurrences and does not anonymize data.

4. Run the impact exercise

The code below is an original local Python model with no cloud calls or real data. It reads JSON containing before, after and consumers. Each field declares name, type, nullable and unit. Types are integer, decimal, string and boolean; type and unit comparisons are exact. Each consumer declares its required fields and allowAdditionalFields. There are no implicit conversions or automatic consumer discovery. In the included example, amount changes from integer/cents to decimal/EUR and riskBand is added. Reader close accepts additional fields but requires the old amount: it fails on type and unit. Reader strict also rejects the additional field. Reader ids only requires id and accepts the rest, so it remains compatible. Reader new-api already expected the new contract and becomes compatible. Reader audit requires checksum, missing both before and after; it remains incompatible without representing a regression introduced by this change. Run python3 content/labs/pde-contract-impact/run.py < content/labs/pde-contract-impact/case.json from the project root. Predict the result before running it. newlyIncompatible should contain close and strict; incompatibleAfter also contains audit. Explain in writing why these sets differ. Then change only allowAdditionalFields and observe that fixing additional fields does not repair type or unit mismatches.

5. Interpret limits and test assumptions

The model distinguishes four transitions: still compatible, newly incompatible, newly compatible and still incompatible. Results are sorted so that they do not depend on input order. Duplicate fields, unexpected properties and values outside the contract are rejected. Input is bounded to 100 fields per schema, 100 consumers and one million CLI characters. These bounds make the exercise understandable; they do not represent cloud product limits. Try both directions of nullability. A producer that may supply null does not satisfy a reader that forbids it; a non-nullable producer satisfies this part of a reader contract that accepts null. Remove a required field and confirm the missing reason. Keep the type but change cents to EUR: the unit remains an independent incompatibility. Exact comparison also treats units with different capitalization as different. The result does not validate actual rows, permissions, allowed values, file encoding or dependency discovery. The actualRowsValidated, allConsumersDiscovered, cloudSchemaValidated and productionApproval indicators remain false. During review, identify the additional evidence needed for each gap. A green report supports further investigation within this contract; it does not authorize a production change.

6. Portability requires behavioral testing

Do not use the exercise rules as an Avro or BigQuery specification. In Avro 1.12.0, writer-to-reader schema resolution has its own rules. A field required by the reader but absent from the writer can use the reader's default. A default on a writer field does not make that field optional during encoding. Always identify which schema contains the default before drawing a conclusion. A shared API also does not guarantee identical capabilities across runners. When moving an Apache Beam pipeline, consult the runner capability matrix for the features used and execute representative cases on the destination. Record versions and restrictions; do not memorize a support table as though it were permanent. If the case depends on timers, state or windows, the test must exercise that dependency. An open format reduces one dependency but does not remove specific SQL, identities, policies or transactional semantics. Build a small matrix of features, evidence and gaps before estimating effort. In a BigQuery script with an exception handler, recording an error does not itself execute ROLLBACK TRANSACTION. Make transaction handling explicit on the failure path and test it in the context being used.

7. Prepare cutover and fallback with data

Replication lag of zero describes a moment. If the source still accepts writes, that value does not establish a stable cutover boundary. Define how to control new writes, drain in-flight work, identify the target catch-up point and confirm application through that point. Coordinate consumers and functional acceptance; changing an endpoint does not resolve movements in transit. Rehearse fallback with the same care. If the target accepts new writes, changing DNS back does not transport them to the source. Identify how those writes would be reconciled and whether a reverse transformation exists. A daily total cannot reconstruct every movement that produced it. Where information has been lost, do not promise lossless fallback simply because an old copy still exists. Monetary conversion also needs an explicit policy. The value EUR 0.015 corresponds to 1.5 cents and cannot be represented exactly as an integer. Rounding is a consequential decision, not proof of reversibility. Define exception handling, precision, reconciliation and the approval owner. In parallel rehearsals, isolate external effects so that two paths do not publish the same real outcome.

8. Hand over the decision and maintenance to RUN

Prepare a short record that allows another team to repeat the reasoning. Include the previous contract, proposal, consumer inventory, new incompatibilities, pre-existing problems, authorized real-data results and rehearsal limitations. Assign each action an owner, deadline and acceptance condition. An old issue can still prevent cutover even when this change did not cause it. In our case, close and strict need a coordinated transition; audit needs separate treatment for checksum. The improvement for new-api does not offset broken close processing. For a consumer that has not yet been inventoried, preserve the uncertainty instead of marking it approved. The committee receives business impact, options and evidence, with concrete criteria for proceeding, postponing or stopping. As a final exercise, write a sequence of five decisions: inventory and fix contracts, check security and compatibility, reconcile data and behavior, rehearse cutover and fallback, confirm acceptance and support. Explain where you would pause if the monthly exporter were missing. Connect this lesson to replay, history and capacity: compatibility defines what can be read; reconciliation confirms what was delivered; operations sustain delivery within the deadline.

"""Original offline field-contract model, not an Avro/BigQuery schema validator."""
import json
import re
import sys


def identifier(value):
 if not isinstance(value, str) or not re.fullmatch(r'[A-Za-z0-9_-]{1,64}', value):
 raise ValueError('invalid identifier')


def fields(rows):
 if not isinstance(rows, list) or not 1 <= len(rows) <= 100:
 raise ValueError('fields must contain 1..100 entries')
 result = {}
 for row in rows:
 if not isinstance(row, dict) or set(row)!= {'name', 'type', 'nullable', 'unit'}:
 raise ValueError('invalid field shape')
 identifier(row['name'])
 identifier(row['unit'])
 if row['name'] in result:
 raise ValueError('duplicate field')
 if row['type'] not in ['integer', 'decimal', 'string', 'boolean'] or type(row['nullable']) is not bool:
 raise ValueError('invalid type or nullability')
 result[row['name']] = row
 return result


def check(producer, reader, additional):
 issues = []
 for name in sorted(reader):
 expected = reader[name]
 if name not in producer:
 issues.append({'field': name, 'reason': 'missing'})
 continue
 actual = producer[name]
 for property_name in ['type', 'unit']:
 if actual[property_name]!= expected[property_name]:
 issues.append({'field': name, 'reason': property_name,
 'expected': expected[property_name], 'actual': actual[property_name]})
 if actual['nullable'] and not expected['nullable']:
 issues.append({'field': name, 'reason': 'nullable', 'expected': False, 'actual': True})
 if not additional:
 for name in sorted(set(producer) - set(reader)):
 issues.append({'field': name, 'reason': 'unexpected'})
 return issues


def evaluate(data):
 if not isinstance(data, dict) or set(data)!= {'before', 'after', 'consumers'}:
 raise ValueError('invalid input shape')
 before, after = fields(data['before']), fields(data['after'])
 consumers = data['consumers']
 if not isinstance(consumers, list) or not 1 <= len(consumers) <= 100:
 raise ValueError('consumers must contain 1..100 entries')
 seen, results = set, []
 for consumer in consumers:
 if not isinstance(consumer, dict) or set(consumer)!= {'id', 'fields', 'allowAdditionalFields'}:
 raise ValueError('invalid consumer shape')
 identifier(consumer['id'])
 if consumer['id'] in seen:
 raise ValueError('duplicate consumer')
 seen.add(consumer['id'])
 if type(consumer['allowAdditionalFields']) is not bool:
 raise ValueError('additional-field policy must be boolean')
 reader = fields(consumer['fields'])
 old = check(before, reader, consumer['allowAdditionalFields'])
 new = check(after, reader, consumer['allowAdditionalFields'])
 transition = ('still-incompatible' if old else 'became-incompatible') if new else (
 'became-compatible' if old else 'still-compatible')
 results.append({'consumer': consumer['id'], 'beforeIssues': old,
 'afterIssues': new, 'transition': transition})
 results.sort(key=lambda row: row['consumer'])
 return {'results': results,
 'newlyIncompatible': [r['consumer'] for r in results if r['transition'] == 'became-incompatible'],
 'incompatibleAfter': [r['consumer'] for r in results if r['afterIssues']],
 'actualRowsValidated': False, 'allConsumersDiscovered': False,
 'cloudSchemaValidated': False, 'productionApproval': False}


def unique_object(pairs):
 result = {}
 for key, value in pairs:
 if key in result:
 raise ValueError('duplicate JSON field')
 result[key] = value
 return result


if __name__ == '__main__':
 try:
 raw = sys.stdin.read(1000001)
 if len(raw) > 1000000:
 raise ValueError('input too large')
 print(json.dumps(evaluate(json.loads(raw, object_pairs_hook=unique_object)), sort_keys=True))
 except (ValueError, TypeError, RecursionError) as error:
 print('invalid input: ' + str(error), file=sys.stderr)
 sys.exit(2)
IN PRACTICE

close and strict break after type and unit changes; audit already required a missing checksum.

Common pitfalls

Confuse regressions with earlier problems, schema with fidelity, dry run with enforcement and endpoint reversal with data recovery.

Related topics: Replay and ingestion · Reconciliation and grain · Continuity and recovery

Take this idea with you

Compare before and after per consumer; validate data and behavior before approving cutover.

Create account

Reference: Professional Data Engineer standard exam guide · Current linked standard guide (document title v4.2); edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.