Cloud Incident Response on AWS: Where to Look First
Blue-team guide to AWS incident response: the log sources that matter, detection signals, and evidence-safe containment.
In this article
When an alert fires in an AWS account, the difference between a two-hour investigation and a two-week one is knowing where to look first. Cloud incident response is not the same as investigating a laptop or an on-premises server: there is no disk to image in the traditional sense, the control plane is an API, and the most important evidence is often a log entry that expires if you did not turn logging on in advance. This guide is written for defenders. The goal is to understand the attack surface so you can detect abuse and recover safely, not to teach anyone how to break in. We will walk through the AWS telemetry that matters, the signals that separate noise from a real intrusion, and how to contain damage without destroying the very evidence you need.
Why cloud incident response is different#
In a traditional environment you respond to a host: you isolate it, capture memory, and image the disk. In AWS the primary object of investigation is the account and its identities. An attacker who obtains a valid credential can act from anywhere in the world through the same API you use, and their actions look syntactically identical to legitimate administration. There is no malware to find on a control-plane compromise, only a sequence of API calls. This shifts the center of gravity of the investigation from binaries and files to identity, entitlements, and audit logs. It also means that preparation is decisive: the evidence you can collect after the fact is exactly the evidence you configured before the incident. If organization-wide logging was off, the reconstruction you can perform is fundamentally limited.
Prepare before the incident: readiness that pays off#
Readiness is the cheapest control you will ever buy. Before anything happens, confirm that a multi-region, organization-level CloudTrail trail is enabled and delivering to a dedicated, write-once logging account that responders can read but service accounts cannot delete. Enable GuardDuty across all regions and all accounts, turn on Config to record resource state over time, and make sure VPC Flow Logs and relevant service logs (S3 access logging, load balancer logs, DNS query logging via Route 53 Resolver) are captured centrally. Establish break-glass roles with strong MFA and document who may assume them. A short, rehearsed runbook that names the log locations, the isolation steps, and the escalation contacts will save more time during a real event than any tool.
The core AWS logging sources you will lean on#
Four sources carry most cloud investigations. CloudTrail is the audit log of the control plane: every API call, who made it, from which IP and identity, and whether it succeeded. GuardDuty is a managed detector that correlates CloudTrail, DNS, and flow data into findings such as credential exfiltration or anomalous API usage. VPC Flow Logs record network conversations at the ENI level, which is where you reconstruct lateral movement and data egress. Config gives you a timeline of how a resource looked at each point in time, invaluable for proving when a security group was opened or a policy was changed. Around these sit service-specific logs (S3, RDS, Lambda, CloudFront, WAF) that add context for the workloads involved.
Where to look first: CloudTrail#
CloudTrail is almost always the first stop. Start by pivoting on the identity in the alert and asking three questions: what did this principal do, from where, and when did the behavior change. Filter for eventName values that indicate reconnaissance or escalation attempts, and pay special attention to management-plane events. High-signal events for a compromise include changes to identity and trust, such as CreateUser, CreateAccessKey, AttachUserPolicy, PutUserPolicy, UpdateAssumeRolePolicy, and CreateLoginProfile. Watch for ConsoleLogin records without MFA, for GetCallerIdentity immediately followed by broad listing calls, and for spikes of AccessDenied results that reveal an actor probing for permissions. Correlate the source IP, the user agent, and whether the calls came through temporary role credentials, which tells you whether an access key or an assumed role was abused.
Identity and access signals#
Because identity is the battlefield, invest in the entitlement view. Enumerate recently created or modified IAM users, roles, and access keys, and compare them against your change records. A newly minted access key on a service account, a role trust policy that suddenly allows an external account, or an inline policy granting iam:* are strong indicators of an attacker establishing durable access. GuardDuty findings in the UnauthorizedAccess, CredentialAccess, and Persistence families are worth triaging immediately. IAM Access Analyzer helps you spot resources that became reachable from outside the account. Throughout, remember that legitimate automation also creates keys and roles, so the signal is the deviation from baseline, not the action in isolation.
Network and data-plane telemetry#
Once you understand what the identity did, follow the data. VPC Flow Logs let you see which instances talked to which endpoints and how many bytes moved, so you can distinguish routine traffic from bulk egress to an unfamiliar destination. DNS query logs frequently reveal command-and-control or staging domains that raw IP flow data hides behind CDNs. On the storage side, S3 server access logs and CloudTrail data events for S3 show object-level reads: a sudden series of GetObject calls across an entire bucket, or a ListBuckets followed by targeted downloads, is the classic shape of exfiltration. For databases, review RDS logs and any query auditing you enabled. The question you are answering is simple and consequential: did data leave, and if so, what and how much.
Turning signals into detections#
Effective hunting means writing the questions down as repeatable queries. Using CloudTrail in Athena or your SIEM, build detections for console logins from new countries or ASNs, for API calls whose user agent does not match your known tooling, for any use of the root account, and for disabling of security services such as StopLogging, DeleteTrail, DeleteFlowLogs, or DisableSecurityHub. Alert on first-seen pairings of principal and region, because attackers often operate in regions your teams never use. Rank findings by blast radius: a change to an identity that can assume other roles matters more than a single denied call. Tune aggressively so that the alerts that page a human are the ones that genuinely warrant waking someone up.
Containment without destroying evidence#
Containment in the cloud is fast, which is a gift and a hazard. You can revoke a session, deactivate an access key, or detach a policy in seconds, but if you delete the compromised role or terminate the instance you may erase evidence you have not yet collected. The evidence-safe sequence is to first snapshot and preserve: take EBS snapshots of affected volumes, capture instance metadata, and export the relevant log ranges to your case store. Then constrain the identity by attaching an explicit deny or revoking sessions rather than deleting the principal, and isolate compute by moving it to a quarantine security group with no egress. Rotate exposed credentials and invalidate temporary tokens. Only after preservation and containment do you move to eradication and rebuild from known-good infrastructure as code.
Common pitfalls#
Several mistakes recur. Teams discover during the incident that CloudTrail was single-region or was never delivering to an isolated account, so the attacker's actions in another region are invisible. Responders terminate instances for speed and lose memory and disk artifacts. Investigators trust the timestamps of a single log without correlating clocks across sources. People chase one malicious API call and miss the persistence the attacker planted three steps earlier, such as a lingering access key or a modified role trust. And a frequent unforced error is fixing the symptom, then failing to rotate every credential the actor could have touched, which invites immediate re-entry. Write these lessons into the runbook so the next responder does not relearn them under pressure.
Blue-team checklist#
Use this as a starting sequence. 1. Confirm scope: which accounts, regions, and identities are implicated. 2. Preserve first: export CloudTrail ranges, snapshot volumes, capture instance metadata. 3. Reconstruct the identity timeline from CloudTrail, focusing on IAM and trust changes. 4. Pull GuardDuty findings and correlate with flow and DNS logs. 5. Determine data impact from S3 data events and egress volume. 6. Contain with deny policies, session revocation, and quarantine security groups, not deletion. 7. Rotate every potentially exposed credential and invalidate temporary tokens. 8. Eradicate and rebuild from infrastructure as code. 9. Document the timeline and feed detections back into your monitoring.
FAQ: How long are AWS logs kept by default?#
It depends on the service and your configuration. The CloudTrail console event history is retained for 90 days, but that is not a substitute for a durable trail delivering to S3, which you control and should retain according to your policy. GuardDuty findings, VPC Flow Logs, and service logs each have their own retention that you set. This is exactly why readiness matters: if you rely on default console history, an investigation that starts late may find the earliest evidence already gone. Configure long-lived, tamper-resistant log delivery before you need it.
FAQ: Do I need to image an EC2 instance, or is that obsolete in the cloud?#
Volatile and disk artifacts still matter when a workload is compromised, so imaging is not obsolete for host-level intrusions. The evidence-safe approach is to take an EBS snapshot of the affected volume and, where feasible, capture memory before you stop the instance, then attach the snapshot to a forensic workstation for offline analysis. What changes in the cloud is that for a pure control-plane compromise there may be no host to image at all; the investigation lives entirely in CloudTrail and identity records. Match the collection to the nature of the incident rather than applying one habit everywhere.
Conclusion#
Cloud incident response rewards preparation and discipline. The accounts that recover quickly are the ones that enabled organization-wide, tamper-resistant logging before an incident, that know CloudTrail, GuardDuty, VPC Flow Logs, and Config are the four pillars, and that contain threats without incinerating evidence. Treat identity as the primary battlefield, follow the data to answer whether anything left, and always preserve before you act. Build these steps into a rehearsed runbook, measure how fast you can reconstruct an identity timeline, and keep feeding what you learn back into detections. Done well, cloud response turns a frightening alert into a methodical, repeatable procedure.
