Start with the request path and measurement
At a fictional fund-reporting portal, average latency improves after caching, but less frequently requested reports remain slow. Before adding resources, separate operations, sizes, customers, and periods. Relate user-observed latency, percentiles, errors, throughput, and component waiting time. A fast average can hide a slow tail; p95 is not the average of the fastest 95 percent. Keep periods and populations comparable between versions. Data aggregated only as sum and count normally cannot reconstruct percentiles; retain raw samples or suitable histograms for the measurement mechanism. Use traces and dependency metrics to locate waiting rather than assigning database delay to the web tier. A performance decision should identify the affected operation and a measurable acceptance threshold.
Define representation identity
The cache key determines which requests may share a response. In CloudFront, a cache policy selects headers, cookies, and query strings included in that key; those values also reach the origin. An origin request policy can forward additional metadata without adding it to the key. This is useful for telemetry but dangerous if the value changes content, language, or authorized scope: forwarding a customer identifier does not by itself isolate cached responses. Establish semantic equivalence first, then remove unnecessary high-cardinality dimensions. Disabling caching for a personalized path can be appropriate. Do not confuse cache identity with authorization: separating variants does not replace access checks for each resource. Test with identities that should receive different representations.
Control freshness and file publication
TTL is a freshness policy, not a guarantee of residence until the last second. CloudFront can evict infrequently used objects before expiry. A minimum TTL above zero can retain content despite no-store or private; for a path that must not be stored, check the effective policy rather than origin headers alone. Distinguish client max-age from shared-cache s-maxage while respecting cache-policy limits. Stale content may suit an informational document but not a financial value requiring freshness. Invalidating CloudFront does not automatically empty browser or corporate-proxy caches. Versioned file names help releases and rollback when the manifest points to the intended revision and required older files remain available. Requesting a fresh browser reload is not proof that every cache layer has been refreshed.
Know each mechanism’s scope
Cache behaviors are evaluated in matching order, not by automatically selecting whichever rule looks most specific. Review public and private paths together. Forwarding POST does not mean its response will be cached or that origin failover will repeat that POST at another origin. Design write continuity at the appropriate layer and preserve decisions about retries and business effects. Origin Shield adds a layer that can reduce redundant origin requests, but it does not repair a wrong cache key or create unlimited capacity. Measure the result under representative demand. A high hit ratio establishes reuse only; it does not establish that the response belongs to the correct customer or meets freshness requirements. Include method and path coverage in the acceptance record.
Prepare for misses, updates, and an empty cache
With lazy loading, a miss takes the application to authoritative storage and then populates the cache. The copy can become stale if a write changes only the database. Write-through adds cache updates to the write path, but partial failures still need handling and a new node is not automatically populated with every historical record. Combining mechanisms requires an explicit contract rather than a generic strong-consistency promise. TTL limits a copy’s lifetime without guaranteeing it never becomes stale. Rehearsing an empty cache measures pressure on the origin; a global flush can abruptly increase load. Plan concurrency controls and gradual recovery while retaining correctness and deadline criteria for critical operations. A successfully rebuilt cache is useful only if the source survived rebuilding.
Demonstrate improvement with a bounded model
In the original exercise, 2000 reads arrive each second and each miss creates exactly one origin query. At 95 percent hits, that means 100 queries per second; at 80 percent, it means 400. The model excludes coalescing, retries, writes, revalidation, and AWS limits. It exposes a dependency hidden by an average. A real rehearsal should include the operation mix, relevant datasets, warm and cold caches, and peak conditions, using synthetic or sanitized data. Retain a baseline and compare against limits defined before testing. The PM brings application teams, APS, and data owners together to accept improvement with evidence of correctness, degraded capacity, and observability as well as latency. Record assumptions so later traffic changes can trigger a meaningful review.
read_rate = 2000
hit_ratios = [0.95, 0.80]
origin_queries = [round(read_rate * (1 - hit)) for hit in hit_ratios]
increase_factor = origin_queries[1] / origin_queries[0]
# Original simplified model: one query per miss, no coalescing or retries.
# Results: [100, 400] queries/s, fourfold increase. Not an AWS quota or load test.In a fictional rehearsal, a portal has 98% hits but serves another customer’s report because the origin receives ClientId while the cache key does not distinguish it. The team contains exposure and reviews the contract before optimizing hit ratio.
Common pitfalls
Confusing forwarded metadata with cache identity; trusting no-store alone; accepting average latency or hit ratio as correctness evidence; flushing everything without measuring the origin.
Related topics: Modernization and orchestration
Caching improves service only when it returns the correct representation within the freshness contract and the origin can handle misses.
Reference: CloudFront cache key · SAP-C02