← AWS DevOps Engineer Professional: operations and delivery
21 / 24 · 80 MIN

Log investigation and evidence coverage

Learn to frame an investigation that states sources, interval, query state, and evidence limits.

Start with the operational question

On a fictional instruction-processing platform, the business asks whether failures occurred between 09:10 and 09:25. The team has a central dashboard, but that does not define observed scope. Write down the hypothesis, accounts, Regions, services, and relevant versions before interpreting results. Distinguish error events, requests, and business operations; several rows may belong to one operation. Also record the clock and time field used. A query returning no rows may mean no matches, an excluded source, or delayed ingestion. The report should state what was queried and which gaps prevent a broader claim. This scope statement lets the next shift reproduce the investigation.

Verify shared sources

OAM supports observation of linked accounts within a Region. The sink belongs to the monitoring account and the link to the source; draw that direction in the service inventory. An account acting as both source and monitor does not automatically relay telemetry from its own sources. Confirm the link and effectively shared data types instead of inferring coverage from dashboard presence. In this lesson’s case, two accounts appear while the third missed onboarding. The team can investigate that account through separate authorized access while correcting configuration. It should identify the exception in reporting and test coverage after the change. A successful query cannot compensate for a missing source.

Control the query and retrieve its result

Starting a query returns an identity that should accompany the investigation. Results in Running state are partial; use bounded waiting and distinguish completion, failure, and timeout. Even with Complete, check pagination and limits before calling the first set all rows. Restrict groups and interval to the operational question and expand them when the hypothesis requires it. Each widget refresh can run another query; choose a useful cadence and explicitly cancel queries no longer needed. Retaining query text, queryId, parameters, and observed result makes the conclusion explainable without relying on a single screenshot. Reducing unnecessary scanning must not become omission of relevant sources.

Preserve data meaning

An index can accelerate an equality search but cannot correct predicate meaning. Use exact comparison when looking for a complete identifier; a substring pattern may include other requests and may not benefit from the same index. Also confirm the account where the indexing policy was created and when indexed ingestion began. After stats, available fields are those produced by aggregation; apply event-level filters before losing that detail. Dedup does not turn rows lacking identity into known unique operations. If missing IDs are excluded, report that exclusion. Optimization is useful only when it preserves the population intended for analysis, including explicit treatment of incomplete records.

Retain usable and protected evidence

The team must investigate without unnecessarily exposing sensitive data. A masking policy acts at ingestion and does not by itself resolve exposure in older events. Unmasked access needs its own assessment; at-rest encryption and query authorization are not the same control. At handover, also consider the key protecting historical logs: disassociating it affects new writes but does not remove existing encrypted data’s dependency. Test representative historical retrieval before accepting disabling. For continuous distribution, do not assume export tasks satisfy a latency measured in minutes. Define destination, access, retention, and delivery evidence according to the operational objective and approved preservation period.

Exercise: bound a conclusion

The local model receives expected sources, observed sources, and fictional query pages. It refuses a conclusion when a source is missing, the query is still running, continuation remains uncollected, or coverage parameters are incomplete. It neither executes Logs Insights nor proves the application emitted every required event. Run the cases, add a third source, and explain why zero errors becomes insufficient. Then write a shift-handover note containing hypothesis, period, sources, result, and residual limitation. Even complete results should be compared with symptoms and functional evidence. A technical conclusion should be precise enough to guide the next decision without promising absolute absence of problems.

# Original local evidence review, not a Logs Insights client or completeness proof.
def review_evidence(expected, observed, pages, scope_recorded):
 if not scope_recorded or not expected:
 return "scope incomplete"
 if expected - observed:
 return "source coverage incomplete"
 if not pages or any(p["status"]!= "Complete" for p in pages):
 return "query not complete"
 if pages[-1]["next_token"] is not None:
 return "results not fully retrieved"
 if any(pages[i]["next_token"]!= pages[i+1]["request_token"]
 for i in range(len(pages)-1)):
 return "page chain incomplete"
 if pages[0]["request_token"] is not None:
 return "first page missing"
 return "recorded scope reviewed; emitter completeness still unproven"

expected = {"account-a/eu-west-1", "account-b/eu-west-1"}
page = {"status": "Complete", "request_token": None, "next_token": None}
assert review_evidence(expected, expected, [page], True).startswith("recorded scope")
assert review_evidence(expected, {"account-a/eu-west-1"}, [page], True).startswith("source coverage")
assert review_evidence(expected, expected, [{**page, "status": "Running"}], True) == "query not complete"
assert review_evidence(expected, expected, [{**page, "next_token": "p2"}], True) == "results not fully retrieved"
pages = [{**page, "next_token": "p2"}, {**page, "request_token": "p2"}]
assert review_evidence(expected, expected, pages, True).startswith("recorded scope")
assert review_evidence(expected, expected, [pages[0], {**page, "request_token": "p3"}], True) == "page chain incomplete"
assert review_evidence(expected, expected, [page], False) == "scope incomplete"
IN PRACTICE

Fictional example: the central query finishes without errors in accounts A and B, but the service also runs in C. The result is valid for A and B; the service-wide conclusion remains pending.

Common pitfalls

Common mistakes: assuming transitive sharing, treating Running as final, ignoring pagination, counting unidentified rows as unique requests, and disabling a still-required key.

Related topics: Recovery workflows and preconditions

Take this idea with you

A strong conclusion connects the operational question to observed sources, complete results, and remaining limits.

Create account

Reference: DOP-C02 monitoring and logging objectives · DOP-C02

AWS is a trademark of Amazon.com, Inc. or its affiliates. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by AWS. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.