1. Define useful work before optimizing
A fictional funds application finishes a batch with lower cloud spending. Its report also highlights more attempts per second. Before presenting an optimization, ask how many distinct operations produced valid outcomes within the agreed deadline. Attempts can include retries, failures and repeats. A configuration may look cheaper because it leaves work unfinished. Define the unit with business, application and FinOps owners before the trial: in this block, an expected operation counts only when a valid outcome exists by deadline. Record the expected set and cost period. Use the same scope for both configurations, including shared components allocated to the service by the agreement. Explain whether a value is observed, estimated or still partial. The exercise definition does not represent a procedure from a specific bank; it makes the decision verifiable. The project manager coordinates criteria and owners, APS confirms operational behavior, and FinOps helps ensure financial components are comparable. If a requirement changes, retain that change explicitly. Removing deadline from the denominator merely to improve the indicator changes its meaning and stops answering the original question. Keep the original criterion available so reviewers can distinguish an improvement from a changed measurement contract.
2. Locate the container’s effective limit
Spare node CPU does not mean every container can consume it. On a Linux node, a configured CPU limit can cause throttling when the application tries to exceed it. In the example, the container has 500m, the node has headroom and throttling increases. The next trial should relate the limit to throughput and latency, instead of assuming more disk resolves the restriction. Requests and limits serve different purposes; lowering a request does not automatically raise the CPU ceiling. For memory, write the model’s components before raising concurrency. With a 512 MiB baseline and another 80 MiB per concurrent request, twenty requests require 2,112 MiB. Against a 2,048 MiB budget, the shortfall is 64 MiB. This calculation neither predicts an OOM’s timing nor replaces real measurement; it uses explicit assumptions of equal request cost and no other usage. In a real service, buffers, caches, runtime and workload distribution can change consumption. Testing lower concurrency may protect memory while increasing waiting. Compare effects against requirements, retain headroom and define how to restore the earlier configuration if latency or useful capacity worsens. Record both resource and application observations so the trial can explain the trade-off.
3. Address the dependency preventing progress
More workers do not guarantee more completions. A database may be constrained by connections or work waiting for a transaction. In PostgreSQL 18, wait_event_type=Lock with wait_event=transactionid means waiting for a transaction to finish. Identify the blocking chain and owner before deciding an intervention. Low CPU is compatible with waiting sessions; more CPU does not automatically finish the blocker. Do not turn diagnosis into indiscriminate session cancellation: understand ongoing work and the authorized recovery process. Also calculate the maximum connections an application may open while scaling. In the exercise, the database allows 180 connections, with 30 reserved for other consumers. Each instance can open ten, including overflow. Fifteen instances fit that budget. A pool normally using few connections can reach its maximum under pressure. Include parallel processes, old instances still active during rollout and other applications in operational reasoning. The supplied number is fictional, not a Cloud SQL default limit. Confirm actual configuration and measure acquisition waits, active connections and completions before expanding pools. A larger queue at the database entrance may merely move the bottleneck. Record which limit each proposed change addresses and which dependency still bounds useful progress afterward.
4. Test the load that must be supported
In a closed load-test model, each user waits for an iteration to finish before starting the next. If the application slows, the rate of new iterations can fall. That suits some questions but does not establish behavior under constant external arrivals. When the objective requires that load, consider an open model and confirm what the generator actually managed to start. Desired configuration and observed load can diverge if generation resources are insufficient. Record both, alongside failures and test duration. Control conditions changing between versions. Measuring A with an empty cache and B with already cached requests does not isolate code improvement. Define datasets, operation mix, cache state and warm-up periods. Not every trial should use a warm cache: startup, new data and recovery may require different conditions. Decide which represent relevant requirements and report results separately. Avoid manually correcting numbers with an arbitrary cold-cache discount. If comparison conditions were unequal, repeat the trial and retain the earlier limitation. These exercises discuss test design; they do not execute k6 or measure a real application’s capacity. The handover should make the tested workload reproducible and identify operational conditions still outside the evidence.
5. Calculate cost per outcome with consistent units
A costs 12,000 cents and delivers one thousand distinct valid operations. B costs 9,000 and delivers six hundred. Cost per useful outcome is twelve cents for A and fifteen for B. Total spending fell, but efficiency measured by this unit worsened. Using A’s volume as B’s denominator would produce nine cents and an unsupported conclusion. Keep cost and volume tied to the same plan, period and scope. Every amount in this block is fictional, not a vendor price. Quality and deadline also change the denominator. With 850 valid on-time operations, one hundred valid but late and fifty invalid, the exercise counts only 850. For 17,000 cents, unit cost is twenty. Report the other sets separately to support investigation and recovery. If there is no valid outcome, the division has no defined value. Show spending and zero outcomes instead of displaying zero unit cost. Repeated rows for one ID do not create additional useful units. Counting distinct IDs aids reconciliation but does not prove a financial operation was not executed twice. Keep repetitions flagged for investigation. Explain these distinctions in the review so a lower financial total cannot conceal unmet delivery criteria.
6. Confirm financial scope and residual components
Removing workers does not necessarily remove every cost associated with a service. In the exercise’s contractual model, there are 400 fixed daily units, 200 variable worker units and one hundred retention units. The change removes only the variable 200. Saving is 200 and 500 per day remain. Do not attribute elimination of a component that stays contracted or necessary to the change. If the objective includes decommissioning, explicitly identify each component, dependency, owner and evidence that its cost ended. Financial observation timing matters too. Cloud Billing exports data to BigQuery without a latency guarantee; services may report usage at different intervals. Immediately after migration, an absent row does not establish zero cost. Mark the set as partial and plan reconciliation when additional data arrives. If a decision is needed earlier, separate estimated from observed values, making assumptions and uncertainty visible. Do not invent a replacement cost by copying the project’s cheapest service. The laboratory declines to present definitive unit cost when costsComplete is false. This rule makes the exercise’s limitation explicit; it does not automatically discover whether a real export is complete. Assign an owner and review date to provisional figures so they do not silently become final reporting.
7. Review recommendations against critical periods
A machine-type recommendation is evidence for analysis with a window and assumptions. The inspected Compute Engine documentation describes recommendations using the last eight days and limitations of averages that may miss short spikes. If critical monthly processing did not occur during that period, a recommendation to shrink a VM does not establish sufficient capacity for the next close. Seek observations of that processing or trial a representative workload before turning estimated savings into an operational commitment. At review, connect proposal, hypothesis, trial and criterion. For example, reducing capacity may lower nominal cost but must retain batch deadline, valid outcomes and agreed recovery capacity. Define a controlled window, stop signals and a rollback owner. Do not assume every recommendation requires immediate implementation or should be ignored. Use it to formulate a testable hypothesis. Retain trial results and limitations, including conditions not reproduced. APS needs to operate the chosen configuration after the project; FinOps needs to track savings actually observed. A complete decision distinguishes expected benefit, realized benefit and requirements still needing evidence. This makes the handover useful when workload or pricing assumptions later change, because the team can identify which part of the decision needs revisiting.
8. Batch cost and identity laboratory
The Python code receives cost in cents, expected IDs, attempt records and deadline. An ID counts when it belongs to the batch, has an outcome marked valid and completed by deadline, including the boundary. Repeats count once in the denominator and remain flagged. IDs outside the batch are reported separately. Validity is supplied by the exercise rather than calculated by inspecting a real transaction. The program neither repairs financial duplication nor establishes application idempotency. Compare plan-a and plan-b. A makes 1,200 attempts, delivers one thousand useful outcomes and costs fifteen cents per outcome. B makes 1,500 attempts, delivers six hundred and costs twenty. Inspect duplicate-row, zero-useful, late-or-invalid and incomplete-cost. Deadline boundary is included and partial cost does not produce a definitive unit value. There are eight fixtures, 4,096 identity and repetition combinations and ten checked invalid inputs. Calculation uses Fraction to preserve exact ratios and leaves inputs unchanged. optimizationApproved remains false: the report supports discussion, not change authorization. Run locally, alter a deadline or outcome and explain why the denominator changed. There are no cloud calls, billing API calls or actual prices. Use the evidence to explain why more attempts can coexist with worse delivery and higher useful unit cost.
"""Original closed-batch accounting exercise, not billing or transaction software."""
from copy import deepcopy
from fractions import Fraction
import hashlib
import itertools
import json
from pathlib import Path
def integer(value, label):
if type(value) is not int or value < 0:
raise ValueError(label + ' must be a nonnegative integer')
return value
def summarize(cost_cents, expected, attempts, deadline, costs_complete=True):
integer(cost_cents, 'cost_cents'); integer(deadline, 'deadline')
if type(costs_complete) is not bool:
raise ValueError('costs_complete must be boolean')
if not isinstance(expected, list) or not expected or any(type(x) is not str or not x for x in expected) or len(set(expected))!= len(expected):
raise ValueError('expected IDs must be unique nonempty strings')
if not isinstance(attempts, list):
raise ValueError('attempts must be a list')
wanted, useful, outside, occurrences = set(expected), set, set, {}
for row in attempts:
if not isinstance(row, dict) or set(row)!= {'operation', 'valid', 'completedAt'}:
raise ValueError('invalid attempt shape')
if type(row['operation']) is not str or not row['operation'] or type(row['valid']) is not bool:
raise ValueError('invalid attempt identity or validity')
integer(row['completedAt'], 'completedAt')
key = row['operation']
if key not in wanted:
outside.add(key)
elif row['valid'] and row['completedAt'] <= deadline:
useful.add(key)
occurrences[key] = occurrences.get(key, 0) + 1
return {'recordedCostCents': cost_cents, 'costsComplete': costs_complete,
'attemptRows': len(attempts), 'usefulCount': len(useful),
'costPerUsefulCents': str(Fraction(cost_cents, len(useful))) if useful and costs_complete else None,
'usefulIds': sorted(useful), 'missingIds': sorted(wanted - useful),
'outsideScopeIds': sorted(outside),
'repeatedUsefulIds': sorted(k for k, n in occurrences.items if n > 1),
'meetsRequiredVolume': useful == wanted, 'optimizationApproved': False}
def row(key, valid=True, time=100):
return {'operation': key, 'valid': valid, 'completedAt': time}
def main:
expected = ['op-' + str(n) for n in range(1000)]
a = [row(k) for k in expected] + [row(expected[i], False) for i in range(200)]
b = [row(k) for k in expected[:600]] + [row(expected[i % 600], False) for i in range(900)]
definitions = [
('plan-a', 15000, expected, a, 100, True, 1000, '15'),
('plan-b', 12000, expected, b, 100, True, 600, '20'),
('duplicate-row', 120, ['a', 'b'], [row('a'), row('a'), row('b')], 100, True, 2, '60'),
('zero-useful', 5000, ['a'], [row('a', False)], 100, True, 0, None),
('late-or-invalid', 200, ['a', 'b', 'c'], [row('a'), row('b', True, 101), row('c', False)], 100, True, 1, '200'),
('outside-scope', 200, ['a', 'b'], [row('a'), row('x')], 100, True, 1, '200'),
('incomplete-cost', 100, ['a'], [row('a')], 100, False, 1, None),
('deadline-boundary', 123, ['a'], [row('a', True, 100)], 100, True, 1, '123'),
]
fixtures = []
for name, cost, wanted, attempts, deadline, complete, count, unit in definitions:
before = deepcopy((wanted, attempts))
result = summarize(cost, wanted, attempts, deadline, complete)
assert result['usefulCount'] == count and result['costPerUsefulCents'] == unit
assert (wanted, attempts) == before
# Keep evidence compact; full IDs remain available from summarize.
compact = {k: v for k, v in result.items if k not in ['usefulIds', 'missingIds']}
compact['missingCount'] = len(result['missingIds'])
fixtures.append({'id': name, **compact})
checks = 0
small = ['a', 'b', 'c', 'd', 'e', 'f']
for mask, duplicates in itertools.product(range(64), repeat=2):
attempts = [row(k) for i, k in enumerate(small) if mask & (1 << i)]
attempts += [row(k) for i, k in enumerate(small) if mask & duplicates & (1 << i)]
result = summarize(120, small, attempts, 100)
count = mask.bit_count
assert result['usefulCount'] == count
assert result['costPerUsefulCents'] == (str(Fraction(120, count)) if count else None)
assert len(result['repeatedUsefulIds']) == (mask & duplicates).bit_count
assert len(result['missingIds']) == 6 - count
checks += 1
invalid = [(-1, ['a'], [], 100), (True, ['a'], [], 100), (1, [], [], 100),
(1, ['a', 'a'], [], 100), (1, [''], [], 100), (1, ['a'], {}, 100),
(1, ['a'], [{}], 100), (1, ['a'], [row('a', 1)], 100),
(1, ['a'], [row('a', True, -1)], 100), (1, ['a'], [], 1.5)]
for args in invalid:
try:
summarize(*args)
except ValueError:
pass
else:
raise AssertionError('invalid input accepted')
print(json.dumps({'scriptSha256': hashlib.sha256(Path(__file__).read_bytes).hexdigest,
'fixtures': fixtures, 'identityCombinations': checks, 'invalidInputs': len(invalid),
'inputPreserved': True, 'cloudExecuted': False, 'billingApiCalled': False,
'network': False, 'persistentWrites': False,
'limitations': 'Fictional complete-scope costs and supplied validity; repeated rows flagged, no proof or repair of financial effects, no optimization authorization.'}, indent=2))
if __name__ == '__main__':
main
A lower-spending plan can move from 15 to 20 cents per useful outcome and fail the required volume.
Common pitfalls
Counting retries as outcomes, ignoring container limits, multiplying pools without a budget or treating incomplete recent exports as zero cost.
Related topics: Queue recovery and reconciliation · Promotion signals and comparable populations · Commitments, reservations and cost scope
An optimization needs to improve outcomes within requirements, using comparable costs and demonstrated operational limits.
Reference: Professional Cloud DevOps Engineer exam guide · Current linked guide; edition date unconfirmed (2026-09-30 inspection)