1. Restrict action and understand impact
A suspicious session may remain valid while applicable permissions change. Authorization is evaluated per request; token validity does not freeze original Allows. Confirm propagation and observed effects on the relevant path. If a role is shared, a broad denial can interrupt healthy consumers. Inspect attached entities before editing a shared customer managed policy. An aws:userid selector targeting a session reduces scope, but the name can change at a later role hop. Investigation must map derived sessions and resource grants. Do not close containment merely because you removed an Allow from an identity policy when another applicable access path exists.
2. Preserve within a system that replaces resources
Production seeks to repair unhealthy instances; investigation may need to preserve precisely one of those instances. DisableApiTermination protects against certain EC2 API terminations but does not prevent Auto Scaling from terminating the resource. Scale-in protection also does not cover health-check replacement. Identify lifecycle mechanisms before choosing the control and coordinate clean service capacity. A temporary measure needs an owner, scope, and exit condition. In a fictional fund close, preserving one node without alternate capacity can harm processing, while using a suspicious image for replacements can spread the compromised condition. The technical manager should make these dependencies visible to the decision without replacing security evidence with schedule pressure.
3. Know what a snapshot contains
An EBS snapshot represents data written to the volume at request time. Transfer may continue while its state is pending; that does not turn capture into a window including every write until completed. Data held only in cache is not guaranteed in that artifact. Balance consistency needs, volatile collection, and containment, recording what each artifact actually covers. Encryption status is inherited from the volume: an unencrypted source produces an unencrypted snapshot, requiring an encrypted copy if that is the requirement. Retain control over both versions and the keys needed for future reads. In the inventory, record request time, completion, volume identity, and known capture limitations.
4. Contain while retaining an investigation path
A deny-all NACL on a subnet has shared scope and can affect healthy nodes. An isolation VPC is an architectural option, but an existing EC2 instance does not simply change VPC while continuing to run. The plan must address preservation, relaunch, and authorized investigation access. If using EC2 Triage, granting SendCommand to the responder does not establish that the target agent is reachable. Test permissions and communication before depending on that path in a crisis. For forensic analysis, protect the original and use a working copy, particularly when a tool can modify metadata. Hashes help verify transfers but do not prove source-system integrity or that everything relevant was collected.
5. Separate approval, request, and outcome
Automation needs to represent three questions: was the action authorized, was the request accepted, and was the expected result observed? An aws:approve gate requires the configured decision; delivered notification or silence until timeout is not approval. Define failure and escalation paths when an approver is unavailable. After requesting an asynchronous operation, use state observation, for example aws:waitForAwsResourceProperty, before advancing. Choose a property, expected value, deadline, and failure behavior matching the criterion. A fixed sleep or returned ID does not establish completion. The local model below keeps these three dimensions separate and treats missing observation as unresolved rather than executing any AWS action.
6. Recover and hand over a trusted service
Eradication should address the cause and affected resources. Rebuilding from the same vulnerable image can restore both service and the entry condition. Use a clean, corrected baseline while first preserving required evidence. Recovery requires customer owners to restore and validate applications; AWS Security Incident Response provides guidance rather than directly performing every restore. Define technical and business checks, increased observation, and criteria for removing temporary measures. In the exercise, calculate the expected result for absent approval, failed request, unknown state, pending, and completed. At real handover, deliver results and exceptions with owners rather than one green indicator. Closure needs to demonstrate the agreed operational objective.
# Original local evidence-gate model; not an SSM executor or incident policy.
# No credentials, network calls, or AWS changes.
def gate(approved, request_accepted, observed_state):
if not approved:
return "approval-required"
if not request_accepted:
return "request-failed"
if observed_state is None:
return "observation-required"
if observed_state == "pending":
return "wait"
if observed_state == "completed":
return "ready"
return "escalate"
cases = [
(False, False, None, "approval-required"),
(True, False, None, "request-failed"),
(True, True, None, "observation-required"),
(True, True, "pending", "wait"),
(True, True, "completed", "ready"),
(True, True, "error", "escalate"),
]
for approved, accepted, state, expected in cases:
assert gate(approved, accepted, state) == expected
print("six evidence-gate cases passed")
Collection was approved and returned an ID, but its state query failed. The artifact remains unconfirmed while the automation owner addresses the missing permission.
Common pitfalls
Frozen permissions; EC2 protection as universal protection; pending as continuous capture; approval as outcome; hash as completeness proof.
Related topics: Response, containment, and preservation · KMS: delegation, context, and recovery
Confirm each stage with matching evidence and keep service impact explicit.
Reference: Disabling permissions for temporary credentials · SCS-C03