Threat Hunting

AI Threat Hunting: How AI Is Changing Security Operations

Learn what AI threat hunting contributes to a hunt, where human analysts still outperform a model, and how to judge any vendor's detection accuracy claim.

Mars Security

Mars Security

Mars Security Research

Your CISO forwarded a vendor deck on Tuesday and asked, politely, whether the hunting team still needs three people. The deck said AI threat hunting on nearly every slide and never once said what the model actually does. Before that conversation goes further, here is the honest version.

The answer, before the detail: a model contributes coverage and continuity to a hunt, not final judgment. It can run the same hypothesis across every host, identity, and cloud account it can reach, every day, and hold a pattern together over a window far longer than anyone can carry in working memory. Picking the hypothesis worth running, and deciding what a shortlist means, stays with the hunter.

Key takeaways about AI threat hunting

  • A model's contribution to AI threat hunting is reach: the same question asked everywhere, every day, across months of history rather than a shift's worth of attention.
  • Responsibility for hypothesis quality stays human. A model can rank a hundred candidates; without the right business context, it cannot tell you that last quarter’s acquisition explains eighty of them.
  • Reported detection accuracy is inseparable from the dataset a model was measured on. Your network is a dataset nobody has published results for.
  • On a fixed detector, lowering false negatives usually raises false positives. Claims that tuning improved both need to be tested on the same dataset, definition, and operating point.
  • Corroboration across separate telemetry sources often beats a higher anomaly score. It is what turns an anomaly into something worth an analyst’s hour.

What Is AI Threat Hunting and What Can AI Actually Do?

AI threat hunting applies statistical models and language models to the search for attacker behavior, including activity that no rule has fired on yet. Strip the packaging and four jobs are left, and all four are worth having.

  • Reach across assets. One hypothesis, tested against every endpoint, identity, and cloud account you own, rather than the twelve a hunter had time for.
  • Reach across time. Sequences held together across sixty or ninety days, where a human reconstructs a week and a threshold rule sees only the current interval.
  • Translation. The same behavioral question expressed against a SIEM, an EDR, an identity provider, and a warehouse, without a person rewriting it four times.
  • Ordering. Ten thousand candidates ranked so the first fifty are worth reading. What no model can own is the hypothesis. Something has to decide that a service account used interactively at 03:00 is worth asking about, and that decision comes from knowing what the account is for. Models generate variations on a question.

So the contribution isn't intelligence. It's coverage and continuity.

AI vs. Human-Led Hunting: Where Each Performs Best

Put AI threat hunting vs manual hunting side by side and the split falls along two axes: breadth and context.

The model does this betterThe hunter does this betterWhy
Testing one hypothesis across every asset, dailyChoosing which hypothesis is worth testing this monthHypothesis quality comes from threat knowledge and business knowledge; only one is in your telemetry
Holding a sequence together across sixty daysReading three events and recognizing a familiar operatorPattern memory over long windows is mechanical; recognition of tradecraft is experiential
Normalizing schemas across four data storesKnowing which of those stores is lyingField mapping is deterministic; knowing that the EDR stopped reporting on a subnet in March is institutional
Ranking ten thousand candidatesDiscarding eighty of the top hundred in four minutesRanking is a scoring problem; exoneration is a context problem

An experienced hunter carries a model of the business that no telemetry contains. The company acquired a firm last quarter and its domain controllers still replicate on the old schedule. One contractor genuinely does work at 04:00, from a different country, every week. There is an application nobody documented that authenticates as a service account and touches file shares at month end. Each of those produces behavior a model will rank highly and be wrong about. A hunter clears them in minutes because the hunter was in the room when the acquisition closed.

Turn that context into a static suppression list and it ages badly, and an attacker who lands inside one of those exclusions gets a quiet quarter.

Finding Weak Signals and Attack Patterns Humans Can Easily Miss

Take a beacon that phones home once every eleven hours, with jitter, over HTTPS, to a domain that resolves through a legitimate CDN. Sixty days of it. T1071.001, and nothing about any single connection is remarkable.

The behavior. Low-and-slow command and control, paced below any interval a threshold rule catches, spread across a window longer than a hunt cycle.

What the model sees. Inter-arrival times per dest_ip and per JA3 hash across the full retention window, variance in bytes_out that stays suspiciously tight, and the fact that this pairing appeared on exactly one host and never on the other four thousand.

What comes back. A shortlist. Call it forty destinations across the estate carrying the regularity signature the model was asked for.

What the shortlist actually contains. Most of it is a misconfigured update client that lost its randomization, a backup agent on a fixed cadence, a monitoring collector, and a licensing check for engineering software with one very consistent user. Regular, low-volume, encrypted callbacks to a CDN-fronted endpoint are exactly what well-behaved software does. Anyone claiming this signal is high-fidelity on its own has not run it.

The analyst's four minutes per candidate is where the hunt happens: the parent process, whether the binary is signed, whether the account has any business reaching that netblock, and whether it started the week a contractor was onboarded. That is also the work you would do when hunting the agent layer, where a privileged process makes constant outbound calls by design.

A model can watch far more for sixty days. It still cannot tell you which sixty days mattered.

How Mars Security Uses AI to Hunt Across Disconnected Security Tools

The method is less exciting than the phrase suggests.

A behavioral hypothesis gets written once, in terms of behavior rather than in terms of a schema. A model then compiles it into the dialect each store speaks, and the query runs where the data already sits: the SIEM, the EDR, the identity provider, cloud audit trails, Snowflake, Databricks. Results land in Mars ISE for correlation. Nothing is copied first.

Translation is the easy part. Three things are harder:

  • Identity reconciliation. The EDR calls it SVC_BACKUP01, the identity provider calls it a UPN, the cloud audit log calls it a role session name. Until those are the same subject, cross-source correlation produces confident nonsense.
  • Clock skew and ingest lag. Two stores disagreeing by ninety seconds can turn one attacker action into two unrelated events, or fabricate a sequence that never happened.
  • Coverage honesty. A federated hunt returns results from the stores it can reach. A store that has been silently failing for three weeks returns zero findings, which reads exactly like a clean result. Get those three wrong and breadth makes you more confident and less correct. That is the real failure mode of federated hunting.

Using AI to Adapt Hunts as Attacker Behavior Changes

A hunt written in March describes what was known in March. When a report lands describing the same operator moving from scheduled tasks to service-based persistence, the hunt goes stale in a way nobody notices, because it keeps returning zero.

Adaptation means regenerating the hunt's logic from the new behavioral description and re-running it against retained history, not just forward. That second half matters. If the tradecraft changed in April and you only run the updated hunt from today, you have confirmed nothing about April. It is the same symmetry argument as meeting an LLM with an LLM: the side that iterates faster gains an advantage, and iteration on the defensive side means rewriting and re-running, not just adding another rule.

Two limits, stated plainly. A regenerated hunt cannot see past your retention, so a sixty-day question against thirty days of logs is a thirty-day answer wearing a longer label. And every regeneration inherits the previous version's benign matches: the service-persistence variant above will surface your patch management tooling and your endpoint agent's updater on every run, until someone reads them once and writes down why.

Can AI Reduce False Negatives Without Creating More Alert Noise?

Not simultaneously, and this is the part vendors skip.

False negatives and false positives often sit on the same dial. Push a model to miss less and it usually flags more. Push it to flag less and it usually misses more. A genuinely better detector may improve both, but claims that tuning alone did so should be measured on the same dataset, definitions, and operating point.

The published numbers make the same point from a different direction. A 2026 evaluation in Scientific Reports trained a deep neural network and a recurrent network under one protocol and measured them on three benchmark sets. False positive rates stayed below 1% on KDDCup99 and NSL-KDD. On the harder UNSW-NB15 set the same architectures ran under 8%. Same architectures, same evaluation protocol, and the reported bound on false positives widened by nearly an order of magnitude because the dataset changed. Apruzzese, Pajola, and Conti reported the same dataset dependence in IEEE Transactions on Network and Service Management in 2022: detectors that scored well on their own data lost substantial detection performance once the traffic came from somewhere else.

Read that before you accept anyone's accuracy figure. The number describes a dataset. Your environment is a dataset nobody has published results for, and its base rate of malicious traffic may be far below the benchmark. That is the condition under which even a sub-1% false positive rate can bury a team.

This is why an assistant that closes false positives faster is solving the wrong problem. It optimizes the end of the pipeline, leaves whatever generated the volume untouched, and never once asks about the behavior nobody alerted on. Faster triage of a bad queue is a faster bad queue.

The lever worth pulling is corroboration. Require support from separate telemetry sources: the network pattern, process lineage, and identity activity.

Each source stays loose. The conjunction tightens.

The cost is real: corroboration can create a false negative. An attacker who appears in only one telemetry source, while the other two are not reporting, will not corroborate and may not surface. That is the trade being made, and it is the honest reason to keep hunters pursuing exploratory hunches on the side.

How Mars Security Scales Hunting Without Scaling the Security Team

That is what we do at Mars Security. We convert threat intelligence into behavioral hunts, run them continuously against the data where it already lives, and hand analysts a corroborated shortlist rather than a scored queue. No new ingestion pipeline. No rip-and-replace. No linear increase in hunting effort. The hunters keep the judgment work. The machine takes the part that was never judgment, which is asking the same question five thousand times.

Frequently Asked Questions About AI Threat Hunting

Can AI replace threat hunters?

Not fully, and the honest reason is not sentiment. A hunt has three parts: forming a hypothesis, executing it broadly, and judging results. Models are strongest at the middle part; the other two depend heavily on knowing what your business does. What changes is the ratio: hunters spend less time querying and more time deciding.

How can AI help with proactive threat hunting?

By removing the sampling. Manual hunting picks a subset of assets and a subset of time because attention is finite, and attackers live in the unsampled remainder. A model can run the full hypothesis across the available estate on a schedule, turning a quarterly exercise into a standing one without scaling human effort at the same rate.

What should I look for in AI-driven threat hunting platforms?

Ask where the query executes, what happens when a data source silently stops reporting, whether hunts can be re-run against history rather than only forward, and how findings are corroborated before an analyst sees them. The category lacks a consistent market definition, so vendor-assigned labels are worth little on their own.

What data does a model need before AI hunting is worth running?

Process lineage from endpoints, authentication events from the identity provider, and network flow or DNS records make a strong starting point. Missing one source limits which correlations you can verify. Retention matters as much as breadth: thirty days of three sources may tell you more than ninety days of one.

How do you measure whether an AI-assisted hunt is working?

Not by alert volume, which measures the threshold rather than the hunt. Track how many hunts produced a shortlist a human actually reviewed, how many candidates survived that review, and how many hunts returned zero for a reason you verified rather than assumed. Silent zeros are the metric most teams never check.

Does this require replacing the SIEM?

No. A federated hunt reads from the SIEM alongside the EDR, identity provider, and cloud logs, and pushes the behavioral question outward rather than pulling the data inward. The SIEM keeps its compliance and retention jobs. What changes is that it stops being the only place a question can be asked.

Give the model the sixty days. Keep the verdict. If that division of labor is the one you want, continuous hunts across your stack is the shape it takes in practice.