Threat Hunting

Cloud Threat Hunting: How to Find Hidden Threats Across the Cloud

Cloud threat hunting starts after the host is gone. Compare log sources and retention defaults across AWS, Azure, and Google Cloud, then scope your hunt.

Mars Security

Mars Security

Mars Security Research

Something odd shows up at 02:14. By the time anyone looks, the container that produced it has exited, the node behind it has scaled down, and the file system is gone. That is the ordinary starting condition for cloud threat hunting, and it is why the practice works differently here than anywhere else.

So the answer to "how do I hunt when the host was destroyed four hours ago" is that you stop hunting the host. You hunt the records the provider wrote on your behalf: control-plane API calls, identity events, workload audit logs, network flows. The evidence window is set by log retention rather than by disk forensics, which means the first question in any cloud hunt is how far back your sources go.

Key takeaways about cloud threat hunting

  • Cloud hunts run against records, not machines. The workload you want to inspect is usually gone, so the surviving control-plane and identity events decide what questions you are still able to ask.
  • Retention is the real scope limit. Provider defaults range from seven days on a free identity tier to 400 days on a Google Cloud _Required bucket, and many teams have never checked their own numbers.
  • Strong cloud hunts often follow a principal, not a machine. Group activity by role session or service account, then look for the same credential appearing from somewhere it has never appeared before.
  • A hunt that ends in a logging change counts as a success. Turning on an audit log type you had switched off is often worth more than the rule the hunt was supposed to produce.

What Is Cloud Threat Hunting?

Cloud threat hunting is the practice of searching cloud telemetry for attacker behavior, including activity no rule has flagged, often working from provider logs rather than live systems. The loop is the one you already run: form a hypothesis, query for it, reach a verdict, and either close it out or turn it into something durable.

What changes is the surface. Three sources carry much of the evidence:

  • **The control plane. **Recorded API calls that created, modified, read, or deleted a resource, along with who made them and from where.
  • Identity. Sign-ins, token issuance, role assumption, and the policy decisions attached to each.
  • The workload. Container and orchestrator audit records, function invocations, and whatever runtime telemetry you chose to collect. Cloud threat hunting isn't endpoint hunting aimed at a new log source. It often centers on identity and the control plane, because that is where cloud intrusions leave some of their most durable marks. An attacker with valid credentials may not need malware, a beacon, or a box you can image. They may need only an API.

Why Cloud Environments Create New Threat Hunting Challenges

Traditional hunting often begins while some of the scene is still there. You get a hostname, pull the process tree, and may still find the answer in memory or on the volume. Cloud hunting begins where that assumption breaks: a crime scene demolished before the investigator arrives.

The demolition is routine and fast. Sysdig's 2025 Cloud-Native Security and Usage Report, published in March 2025, found that 60% of containers live for one minute or less.. The workload that made the suspicious call is not waiting for you.

In the cloud, the host may be gone by the time you have a question about it. The log is often the witness that remains.

Which makes retention the real boundary. Many teams cannot state their own defaults from memory, and those defaults set the outer edge of the hunts they can run. Seven days of sign-in history puts a quarterly pattern out of reach. Ninety days of control-plane events makes longer dwell time unreachable, not undetected.

The container. The credential. The node. The network namespace. Gone.

Hunting Across Cloud Activity, Identity, and Workload Data

Before you write a query, know what each source answers and how long it keeps the answer. Current provider defaults:

Telemetry sourceWhat it recordsWhat it will not tell youDefault retention
AWS CloudTrail management eventsRecorded management events: principal, action, source IP, and user agentData-plane reads unless you enable data events, and nothing inside the guestEvent history holds 90 days; a trail or event data store keeps whatever you configure
Azure Activity LogSubscription-level control-plane operations against resourcesData-plane operations inside a resource, which need a diagnostic setting90 days, then deleted
Google Cloud Audit LogsAdmin Activity and System Event entries, always written and not disableableData Access entries, off by default outside BigQuery_Required bucket 400 days, not configurable; _Default bucket 30 days
Entra ID sign-in logsInteractive and non-interactive sign-ins, plus the conditional access resultWhat the identity did after it authenticated7 days on Free, 30 days on P1 and P2
VPC flow logsAccepted and rejected flows at a network interfacePayload, hostnames, anything inside TLSNone. Flow logs are not created for you, and retention belongs to the destination
EKS control-plane audit logsKubernetes API requests made against the clusterAnything at all until you enable the log type per clusterSet by the CloudWatch log group you route them to

Two rows deserve a second look. Outside BigQuery, Data Access logs on Google Cloud and control-plane audit logs on EKS are both off unless someone turned them on, so a hunt can come back empty for reasons that have nothing to do with the adversary. On GKE, cluster audit records arrive through Cloud Audit Logs instead of a separate switch.

One caution before you widen query access. Logs collect what applications hand them, and secrets have a habit of ending up in telemetry, which makes read access to hunt data a security decision in its own right.

Finding Suspicious Behavior That Traditional Alerts Miss

The obvious objection first, and it is fair. Posture management finds misconfiguration. Cloud-native detection services find known-bad API patterns, and they are good at it: written by the people who built the control plane, updated as the API changes, requiring no pipeline from you. For crypto-mining, known malicious infrastructure, and published abuse patterns, they will often beat what a team writes alone.

The gap is behavioral, not technical. The remaining question is whether an authorized principal did something unusual.

Two cloud threat hunting use cases make it concrete. The first, walked end to end:

A batch job runs in a container and assumes a role through its workload identity. Four minutes later, the container exits and the node scales down. Twenty minutes after that, the same role session is used to enumerate secrets and describe instances, from an address outside your egress ranges.

No malware, no known-bad pattern. The surviving evidence is thin and sufficient: the AssumeRole event with its session name, and every later call carrying that same session with a different sourceIPAddress.

  • The behavior. A role session outliving the workload that requested it, used from a network path the workload never had.
  • The hunt. Group control-plane events by role session, then compare the source address on the assume event against every subsequent call in that session.
  • The benign look-alike. Cross-region service calls, a NAT gateway change, and legitimate session handoffs to a build agent in another account all produce the same shape. Retrieval without triage will bury you. Advanced cloud threat hunting is mostly this: pick a principal, follow it across sources, and look for the seam where its behavior stops matching its purpose.

Connecting Cloud Signals With Endpoint and Identity Evidence

The second scenario is shorter, and it crosses a boundary. A long-lived access key belonging to a CI service account is used at 03:40 on a Sunday to pull an artifact bucket. CI is not running. Nothing in the cloud logs says that is wrong, because the key is valid and the action is one it performs every weekday.

Answering it takes evidence from outside the cloud: which laptop last checked that key out of the secrets manager, whether that endpoint was active at the time, and whether the human attached to it signed in anywhere. The benign look-alike is a scheduled backfill or a pipeline someone re-ran manually after a Friday failure, which is why the endpoint side of the question matters. Endpoint telemetry is a discipline of its own, and this is one place a cloud hunt should reach into it.

Then the honest part. Correlating cloud identity with endpoint identity means reconciling human principals, service accounts, assumed roles, and workload identities that no single system maps back to a person. Many teams do this by hand, with a spreadsheet and institutional memory, and the result is fragile. There is no universal answer, and any tool promising one still depends on mappings and naming conventions you may not have.

What you can avoid is making the seam worse. The industry's reflex when a new surface appears is to sell a new console for it, which leaves you with another identity model to reconcile. The alternative is unglamorous: extend the collectors you already run and keep the joins in one place.

How Mars Security Hunts Across Cloud and Existing Security Data

That is what we do at Mars Security. Cloud logs get queried where they already sit, in the provider's own store, the SIEM, or the warehouse the platform team already pays for, and the same hunt runs against endpoint and identity data without any of it being copied first. No new data copy. No second retention clock. No full normalization project.

The hunts come from threat intelligence mapped to the environment they run in, so they express behaviors rather than indicators: a role session used from an unexpected path, a service account acting outside its window, an audit log type that stopped producing events. The sources stay where they are, and the provider's retention sets the same boundary for us that it sets for you.

Moving From Cloud Hunt Findings to Stronger Detections

Many cloud hunts do not end with an adversary. They end with a gap, and the discipline is in what you do next.

A finding becomes a detection when you can state four things: the behavior, the log source that carries it, the benign population that resembles it, and the retention that supports it. Skip the fourth and you will ship a rule with a seven-day memory against an attacker who works on a quarterly cadence.

Retention shapes the design too. A weekend-activity detection on a CI principal is durable against 90 days of control-plane events and fragile against a free-tier sign-in log. Where the source is too short, write against the longer one and accept the coarser signal, or route the short source somewhere it survives.

  • The behavior. An audit log type that produced events last month and produces none now.
  • The hunt. Compare event counts per log source, per account, week over week.
  • The benign look-alike. A decommissioned account, a workload that stopped running for good reasons, and a sink rewritten during a migration. All three look identical to a source someone switched off. The other honest outcome is a configuration change. Turning on Data Access logs, enabling cluster audit logging, or extending a bucket from 30 days to a year may do more for next quarter's hunts than the rule you were trying to write. Log a finding as a detection, a logging change, or a closed hypothesis. All three are results.

Frequently Asked Questions About Cloud Threat Hunting

What logs do you need to start cloud threat hunting?

Start with control-plane management events and identity sign-ins, since together they cover who did what and how they got in. AWS, Azure, and Google Cloud record core management activity by default, so a first hunt may need no new pipeline. Add workload audit logs and flow logs once you know which accounts and clusters matter most.

How far back can a cloud hunt look?

Only as far as your shortest relevant source. Windows differ by provider and by tier, so a hunt joining sign-ins to control-plane calls is capped by whichever expires first. Audit the numbers before you scope the hunt, because a retention gap looks exactly like an absence of findings.

Can you hunt across more than one cloud without normalizing the schemas first?

Yes, if you map a small set of common concepts rather than fully normalizing every schema. Principal, session identifier, source address, and timestamp appear across provider audit records under different labels. Full normalization is a project; mapping the fields needed for one hunt is a smaller job.

How do you hunt serverless and managed services?

Serverless leaves invocation records and control-plane calls, and little else. You can see that a function ran, which role it used, and what API calls followed. You cannot see what happened inside it unless the application emitted its own telemetry, which makes the identity trail the primary evidence.

How often should a team run cloud hunts?

Frequently enough that the evidence still exists when you get there. If your shortest source holds seven days, a monthly hunt cadence is already looking at a window that has closed. Match the interval to the retention, not to the calendar.

You can't reconstruct the runtime state of a container that exited four hours ago. What you can do is make sure the records that outlived it were switched on, kept long enough to matter, and reachable from one query. Get those three right and the demolished scene stops being a dead end, because the building is gone but the visitor log, the badge reader, and the street camera all kept running. Start by querying cloud telemetry in place instead of moving it somewhere new.