← Professional Cloud DevOps Engineer: delivery and reliability
15 / 15 · 120 MIN

FinOps: performance and useful cost

Connect profiles, billed resources, commitments and valid deliveries to compare production alternatives.

1. Define the value unit before savings

A fictional team prepares a proposal to reduce instruction-processing cost. The useful outcome is a valid instruction completed before its deadline without duplicate effects. Define that unit with the business and preserve it across alternatives. A started attempt, a queued message and a reconciled delivery are different outcomes. If volume or quality changes, explain that change before comparing percentages. The technical manager should be able to connect every cost category to a service and an owner. The exercise uses fictional cost units unrelated to Google Cloud prices. Include compute for every attempt, transfer and telemetry. A real proposal would add other applicable categories, such as storage, licenses, support or migration work. Also state the period and success criterion. An alternative costing less but delivering fewer on-time instructions can fail acceptance. This discipline gives the committee a reproducible comparison and avoids presenting a category reduction as total savings. Retain the baseline and count origins so FINOPS and APS can review the same reasoning.

2. Use profiles to choose the trial

Before buying CPU, identify where requests spend time. In a Java service, high wall time with little CPU time can indicate waiting on a dependency, lock or another condition. Correlate the profile with the slow path. Additional CPU may help an instruction-intensive block but does not demonstrate improvement in an external wait. Define the hypothesis, small change and outcome measurement. Keep version, load and dependency consistent in the trial whenever needed to isolate the effect. In a Go service, allocated heap represents allocations over an interval, including memory already freed. A high total with stable live heap can indicate churn and garbage-collection work; it does not alone prove everything remains retained. Compare profile types supported by the language and preserve the observed window. Avoid using one snapshot to explain the entire business cycle. In the optimization report, state the performance evidence supporting the proposal and the result that would invalidate it. This lesson interprets described profiles; it neither collects real profiles nor executes production workloads.

3. Connect configuration to the billed unit

In a Standard node pool without Autopilot, reducing Pod requests does not automatically change the size or lifetime of VMs still running. It may enable better placement and consolidation, but reduced billed resources must be demonstrated. Check autoscaler limits, Pod movement constraints and remaining capacity before removing nodes. Do not lower requests merely to fit work that then lacks resources. The proposal should include both efficiency evidence and performance and recovery criteria. In Cloud Run, reviewing billing mode requires understanding when the application needs CPU. Mandatory work started in memory after the response is not made safe simply by choosing a billing mode. Request-based billing does not provide an always-allocated CPU assumption; instance-based billing does not turn local memory into durable delivery either. Design asynchronous execution and failure recovery. For sporadic traffic, compare minimum-instance cost with first-request latency after inactivity. A trial using only continuous traffic does not test that requirement. Record the specific execution and billing model instead of generalizing one product’s rules to another.

4. Separate utilization, coverage and eligibility

Define denominators before presenting a commitment. In this lesson’s model, 100 units are bought per hour, 80 of those units are used and total eligible consumption is 160. Commitment utilization is 80/100, or 80%; usage coverage is 80/160, or 50%. One indicator can improve while the other worsens. Neither alone equals net savings, because that comparison needs prices, charges and the alternative without commitment. Also specify whether units represent resources, spend or another measure, avoiding sums of incompatible quantities. A Compute Engine resource-based commitment has region and configuration scope. Moving an application to another region does not automatically carry that coverage, and deleting its VM does not cancel the acquired obligation. Check applicable terms and destination eligibility. Distinguish this mechanism from flexible spend-based commitments. In the project plan, connect purchasing decisions to migration, retirement and demand-stabilization dates. Do not derive a multiyear commitment from one average without understanding service demand and change. This exercise neither recommends nor executes a purchase.

5. Observe each hour and reserved resource

Consider coverage of 100 units in each of two hours. Eligible consumption is 60 in the first and 140 in the second. The average of 100 per hour hides 40 unused commitment units and 40 uncovered usage units. The model has no carryover between hours. Match within each interval before aggregating. The same care applies to migration: a seemingly stable monthly baseline can include periods when workload is ineligible for the commitment being considered. Capacity reservations also need an owner and a review of continued need. A Compute Engine reservation that still exists can incur cost without consuming VMs. That does not imply it should be deleted: it may support an approved recovery requirement. Document the reason, horizon and review criterion. When a VM consumes the reservation, avoid counting the same reserved resource twice in the cost model. If releasing capacity is proposed, coordinate with the continuity-plan owner. Demonstrated savings and recovery capacity should be evaluated in the same operational context.

6. Reconcile exports and additional costs

An export row has cost=120 and credits of −15 and −5. Net cost is 100. If both credits are expanded and cost is summed on each resulting row, the result is 220 because 120 was counted twice. Aggregate credits per original row or use another transformation preserving that cardinality. Also test rows without credits and with several credits. Before presenting real figures, confirm usage or invoice period, currency, scope and included charge types. This arithmetic example does not attempt full cloud-invoice reconciliation. In a separate exercise, moving processing reduces compute by 48 units per day but adds 70 in transfer. With volume, quality and other costs unchanged, the net increase is 22. Evaluate network origin, destination and product under current rules rather than assuming all internal transfer is free. For logs stored in two billable buckets, account for both destinations and retention periods. Matching insertId does not automatically eliminate costs for both copies. Review useful redundancy, investigation and preservation with owners; an exclusion intended to reduce spending should preserve evidence needed by operations.

7. Lab: compare valid deliveries

Save the code as run.py and execute python3 run.py with Python 3.13. The baseline costs 420 in compute, 80 in network and 40 in telemetry, delivering 50 valid on-time outcomes. Total cost is 540 and cost per outcome is 10.8. The incomplete alternative costs 520 but delivers only 40; cost per valid outcome is 13 and it misses the requirement of 50. The candidate costs 485, delivers 50 and yields 9.7 per outcome. Compute costs already include every attempt: do not multiply again by the attempts counter. Predict results before reading the JSON. no_delivery returns null for unit cost because there is no valid denominator. The program checks 61 delivery quantities, 441 hourly-usage pairs and five credit counts, and rejects six invalid inputs. Compare creditNet=100 with naiveExpandedNet=220. Counts are supplied as ground truth and prices are fictional. The program executes neither SQL nor billing APIs nor a statistical retry forecast. A real proposal would still need to demonstrate count origin and quality, measure uncertainty and validate alternative execution.

8. Present the decision and follow the outcome

In the final case, the committee receives a proposal omitting network cost and duplicating costs while expanding credits. Correct both problems before recommending approval. Do not treat the duplication error as hidden contingency: if uncertainty requires a margin, present it separately and explain its basis. Retain assumptions allowing alternatives to be compared and state what still depends on trials. Development can confirm behavior, APS validates operation and FINOPS helps reconcile units and charges. After the change, monitor total cost, valid on-time deliveries, latency and recovery capacity. Compare against the baseline over representative periods and investigate deviations instead of diluting them in an average. Define who decides to continue, correct or revert. Practise a precise business conclusion: “Compute spend is lower, but network charges make the overall proposal more expensive under the current assumptions.” The lesson summary is to define value, measure the actual bottleneck, identify the billed unit, check eligibility and reconcile data before declaring savings. These decisions connect performance, financial management and responsibility for the production service.

"""Original fictional cost comparison. No billing API, SQL engine or purchase."""
from fractions import Fraction
from hashlib import sha256
from pathlib import Path
import json
import platform


def natural(value):
 if type(value) is not int or value < 0:
 raise ValueError('expected a nonnegative integer')
 return value


def assess(compute, network, telemetry, valid_on_time, required, attempts):
 for value in [compute,network,telemetry,valid_on_time,required,attempts]:
 natural(value)
 if valid_on_time > attempts or required == 0:
 raise ValueError('invalid completion or requirement')
 total = compute+network+telemetry # compute already includes every retry.
 return {'total':total,'valid_on_time':valid_on_time,'attempts':attempts,
 'cost_per_valid':str(Fraction(total,valid_on_time)) if valid_on_time else None,
 'requirement_met':valid_on_time >= required}


def hourly(commitment, usage):
 natural(commitment)
 if commitment == 0 or not usage:
 raise ValueError('positive commitment and nonempty window required')
 for value in usage:natural(value)
 used = sum(min(commitment,value) for value in usage)
 bought = commitment*len(usage)
 total = sum(usage)
 return {'used':used,'bought':bought,'usage':total,'unused':bought-used,
 'uncovered':total-used,'utilization':str(Fraction(used,bought)),
 'coverage':str(Fraction(used,total)) if total else None}


def net_row(cost, credits):
 natural(cost)
 if any(type(c) is not int or c > 0 for c in credits):
 raise ValueError('this fixture accepts nonpositive integer credits only')
 return cost+sum(credits)


def run:
 plans={'baseline':assess(420,80,40,50,50,60),
 'incomplete':assess(300,180,40,40,50,75),
 'candidate':assess(360,90,35,50,50,55),
 'no_delivery':assess(20,5,5,0,50,5)}
 assert plans['baseline']['cost_per_valid']=='54/5'
 assert plans['incomplete']['cost_per_valid']=='13' and not plans['incomplete']['requirement_met']
 assert plans['candidate']['cost_per_valid']=='97/10' and plans['candidate']['requirement_met']
 assert plans['no_delivery']['cost_per_valid'] is None
 commitment=hourly(100,[60,140])
 assert commitment==dict(used=160,bought=200,usage=200,unused=40,uncovered=40,utilization='4/5',coverage='4/5')
 assert Fraction(80,100)==Fraction(4,5) and Fraction(80,160)==Fraction(1,2)
 assert net_row(120,[-15,-5])==100
 duplicated=2*120-15-5
 assert duplicated==220 and 70-48==22
 credit_cases=0
 for count in range(5):
 credits=[-3]*count
 assert net_row(120,credits)==120-3*count
 credit_cases+=1
 commitment_cases=0
 for first in range(0,201,10):
 for second in range(0,201,10):
 result=hourly(100,[first,second])
 assert result['used']+result['unused']==200
 assert result['used']+result['uncovered']==first+second
 assert 0<=result['used']<=200
 commitment_cases+=1
 deadline_cases=0
 for valid in range(61):
 result=assess(420,80,40,valid,50,60)
 assert result['requirement_met']==(valid>=50)
 assert (result['cost_per_valid'] is None)==(valid==0)
 deadline_cases+=1
 invalid=0
 for args in [(-1,0,0,1,1,1),(True,0,0,1,1,1),(1.5,0,0,1,1,1),(1,0,0,2,1,1),(1,0,0,1,0,1),(1,0,0,-1,1,1)]:
 try:assess(*args)
 except ValueError:invalid+=1
 assert invalid==6
 return {'scriptSha256':sha256(Path(__file__).read_bytes).hexdigest,
 'pythonVersion':platform.python_version,'plans':plans,'commitment':commitment,
 'creditNet':100,'naiveExpandedNet':duplicated,'migrationNetIncrease':22,
 'creditCases':credit_cases,'commitmentCases':commitment_cases,
 'deadlineCases':deadline_cases,'invalidPlanCases':invalid,
 'network':False,'persistentWrites':False,'vendorExecution':False,
 'independentVerification':False,
 'limitations':'Fictional integer cost units and trusted completion counts. No real tariffs, taxes, full invoice reconciliation, stochastic retry forecast, SQL execution, billing API or resource purchase.'}


if __name__=='__main__':print(json.dumps(run,ensure_ascii=False,indent=2))
IN PRACTICE

Saving 48 compute units while adding 70 network units increases total cost by 22.

Common pitfalls

Confusing requests with billing, coverage with utilization and local savings with total cost, or duplicating costs when expanding credits.

Related topics: SRE and recovery capacity · Observability and evidence preservation · Project management and business case

Take this idea with you

Defensible savings preserve the value unit, include relevant costs and demonstrate the operational result.

Create account

Reference: Profiling concepts · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.