Threat Detection Engineering: How to Automate and Scale
Detection engineering is a function, not a job title. Learn how to measure its rate, find the real bottleneck, automate the right step, and retire stale rules.

Mars Security
Mars Security Research
A report lands in your inbox on a Tuesday. Sixteen pages, one technique worth catching, and a team that will not reach it for five weeks. That distance between reading and detecting is the whole of threat detection engineering, and it is the number nobody on your team is measuring.
Detection engineering is the practice of converting knowledge about attacker behavior into tested detection logic that runs against telemetry you already collect. The output is a deployed, owned, measurable detection. Judge the function the way you would judge any production line: how many behaviors go in, how many working detections come out, and how long the trip takes.
Key takeaways about threat detection engineering
- Threat detection engineering is a production function with a measurable rate. Count behaviors accepted and tested detections deployed. When those two numbers diverge for two quarters, coverage is falling while the backlog grows.
- A common constraint is telemetry rather than talent. Many stalled detections depend on fields the environment does not currently collect, which turns them into a negotiation with a platform owner instead of an engineering task.
- Automating security detection engineering raises rule volume before it raises rule quality. Adopt it with test data drawn from your own environment and a retirement policy written before the first generated rule ships.
- Rule sets decay. A well-maintained public library can spend more of a single release revising and retiring detections than writing new ones, and a private rule set that retires nothing is drifting.
Why Traditional Detection Engineering Can’t Keep Up
Most teams describe the function as a queue. Intelligence arrives, someone reads it, a ticket gets written, and the ticket waits. Counting that queue is easy and tells you almost nothing, because forty open tickets mean one thing at a conversion rate of eight per quarter and something else at thirty. The number that matters is the distance between an observed technique and a tested detection running in production. Authoring speed is a small part of it.
The case for the artisanal version deserves a fair hearing. A detection written by an engineer who understands the environment deeply and tunes it against real traffic can outperform a generic generated rule on precision. The tradeoff is throughput. The problem is arithmetic. That engineer’s output is inherently constrained by the time required for research, testing, and tuning, while new behaviors continue to enter the backlog.
Watch a maintained public rule set and the shape of the work becomes visible. Splunk’s Enterprise Security Content Update v6.5.0, released August 27, 2026, shipped 13 new detections, revised 46 existing ones, and scheduled a 2020-era F5 detection for removal because the platform version it targeted is no longer supported. Even a funded content team can spend most of a release keeping what it already has alive. Your team carries that load too, on top of the backlog, and mostly without counting it. It is a large part of why detection stopped keeping pace.
Turning Threat Intelligence Into Actionable Detection Logic
Take one report and follow it end to end. A published write-up describes an operator registering a scheduled task that runs a signed binary out of a user-writable directory, then moving to a second host over WinRM. Two behaviors, both mappable; neither of them an indicator you can block for longer than a week. This article sits on the conversion side of the intel-to-detection pipeline: one report, one behavior, one deployed detection. Behavior that only becomes visible when four sources are read together is a different problem, and cross-source intelligence gaps cover it.
Here is the calendar a report like that generates.
- Day 0. The analyst on rotation reads it and writes the behavior as one sentence, with the hashes stripped out.
- Day 1. A detection engineer picks the telemetry: process creation with parent lineage, plus the remote management session log.
- Day 4. The engineer finds that Microsoft-Windows-WinRM/Operational is not forwarded from the server estate. The field does not exist.
- Days 4 to 26. Negotiation. The platform owner wants a volume estimate. The estate owner wants a change window.
- Day 26. Logic written against the two fields that do exist, with the third stubbed.
- Day 31. Tested against two weeks of recorded telemetry. Nine hits, eight of them one patch agent.
- Day 34. Deployed, with those eight excluded by service account and path. Thirty-four days, twenty-two of them spent asking somebody for a log. No lifecycle diagram carries that step, and it is where most programs lose their year. Adding another author rarely moves the number, because the constraint sat with the platform owner the whole time.
Mapping Attacker TTPs to Telemetry and Security Controls
Mapping is a feasibility check. Before anyone writes logic, you are asking which behaviors you can afford to detect with the telemetry and controls already in place, and which ones require somebody to change a collection decision first.
| Technique | Telemetry it needs | What already sees it | The gap that stops you |
|---|---|---|---|
| T1053.005 Scheduled Task | Task registration events, process creation | EDR on endpoints | Registration is collected on workstations, not on the server estate |
| T1021.006 Windows Remote Management | Microsoft-Windows-WinRM/Operational, network session records | EDR sees the child process, not the session | Channel not forwarded, so no source-host attribution |
| T1098.001 Additional Cloud Credentials | Identity provider audit log, credential-creation events | Cloud audit trail | Retained 30 days; the campaign you care about is older |
| T1567.002 Exfiltration to Cloud Storage | Proxy or DNS records, plus read volume per session | Proxy, which allows the destination by policy | Volume is the signal, and nobody records bytes per session |
Three of those four gaps are collection decisions taken years ago by someone who has since left. The fourth column is the one worth arguing about in a planning meeting, because it converts an abstract coverage debate into a short list of owners and log sources. The third column matters for the opposite reason: where a preventive control already blocks the behavior reliably, detection may rank lower than an uncovered behavior, although it can still provide attribution, investigation context, and evidence that the control was tested.
Building High-Fidelity Detections Without More Alert Noise
Anchor on the behavior. A scheduled task registered by a process that has no business registering one holds up across campaigns. The binary name, the directory, and the task name will all change by next quarter. Write the version that survives the change.
Name what legitimately looks the same. Backup agents, patch management, and remote monitoring tools register scheduled tasks constantly, from paths that look wrong until you check them. Cloud credential creation has the same problem: pipelines and infrastructure-as-code runs mint service principal secrets on a schedule, and break-glass rotation is indistinguishable from an attacker adding a key. If you cannot name the benign twin, the detection is not finished.
Test against your own recorded telemetry. Lab data tells you the logic parses. Two weeks of your own traffic tells you the hit count, which is the only number that predicts what happens at 02:00 on a Saturday.
Set the retirement trigger when you write the rule. Not a date on a wiki. A condition: the telemetry source disappears, the covered behavior is no longer relevant to the environment, the logic is superseded, or repeated review shows the detection no longer provides useful coverage.
Where this falls apart: automated detection generation can raise rule volume faster than rule quality if generated logic is enabled without production testing. Two guardrails make that stretch survivable. Every generated detection gets tested against real telemetry from your environment before it is enabled, and every detection ships with a retirement trigger attached. A team with no recorded telemetry to test against, or no agreed mechanism for retiring a rule, should build those two things before automating authoring at all.
Automating the Detection Engineering Lifecycle
Automate in the order that moves the rate, and be honest about what each step still needs from a person. Sigma, YARA, and Elastic’s detection-rules repository cover the authoring end well and cost nothing, which is why lean teams start there. What they do not tell you is which candidate is worth authoring, or when to stop running one.
- Intake. Machine reading of reports into candidate behaviors, deduplicated against the rules you already run. Skip the dedupe and you will grow the backlog you meant to shrink.
- Feasibility. Score each candidate against the fields you actually collect. This needs an accurate inventory of collected fields, which most teams do not have and have to build once, by hand.
- Authoring. Generate the logic in your own query language. Highest-variance step in the chain. Treat the output as a first draft with a named human owner, never as a rule.
- Testing. Run the draft against recorded telemetry and count the benign hits before anyone sees an alert. Only as good as the retention window you kept.
- Retirement. Re-evaluate every rule on its trigger and remove what no longer earns its place. Cheapest step to automate, and the one almost everyone skips. A rule set nobody retires is not a rule set. It is a sediment layer. The reason retirement gets skipped is that deleting a detection feels like giving up coverage, when the rule in question has been firing on a decommissioned file server since March.
Continuously Hunting for Threats Existing Rules Miss
Hunting earns its place in a throughput argument for one reason: it produces inputs your rule set cannot generate for itself. A hunt goes looking for behavior no current rule covers, against telemetry you already hold, and it is the only reliable source of candidates that did not come from somebody else’s report.
The limit is plain. Hunting scales with analyst hours, and a hunt that finds nothing still consumes them. What it leaves behind is a dated negative result and a query that ran successfully, which is more than a closed ticket. Both are worth keeping. Only one of them is worth promoting.
Turning Threat Hunts Into Production-Ready Detections
Most hunts die in a chat thread. Somebody finds something interesting on a Thursday, writes it up, and the query is never run again. Promotion is the step that converts hunting effort into throughput, and it needs criteria rather than enthusiasm.
- The query returns a stable field set across the whole estate, not only the three hosts it was written on.
- The benign baseline has been measured and written down as a number.
- Someone owns it by name.
- It carries a retirement trigger. A worked case: a hunt for outbound connections to newly registered domains from server-class hosts. It found two, both explainable. The benign twins are certificate validation traffic, a telemetry vendor’s new endpoints, and marketing tooling that ended up on the wrong VLAN. Once those are baselined, the query is stable enough to run on a schedule, and it becomes a detection.
Not every hunt should be promoted, and forcing it is how sediment forms. A one-time finding tied to a system being decommissioned in six weeks is a ticket. Keep the query, skip the rule.
How Mars Security Automates Threat Detection Engineering
That is the shape of the work, and it is what we built Mars Security to run. Mars reads threat intelligence across the ecosystem, extracts the attacker behavior, and maps it against telemetry you already hold in SIEM, EDR, identity providers, cloud, Snowflake, and Databricks, by querying it where it sits. No ingestion pipeline. No rip-and-replace. The hunt library underneath it was written by people who spent their careers on the offensive side, so it is organized by what an operator has to do. Coverage reads as a live view, and stale detections surface in the same place.
Frequently Asked Questions About Threat Detection Engineering
What is detection engineering in cyber security?
It is the practice of turning knowledge about attacker behavior into tested, deployed detection logic that runs against your telemetry. It is a function rather than a job title, and it usually sits between the intelligence team that reads the reports and the SOC that answers the alerts, reporting to whichever of the two owns coverage.
How is detection engineering different from threat hunting?
Hunting searches for behavior nothing currently detects, and its output is a finding. Detection engineering produces repeatable logic, and its output is a rule someone maintains. The two meet at promotion, when a hunt query stable enough to run unattended becomes a detection. Teams that put both under one lead usually promote faster.
What metrics should a detection engineering team report?
Conversion rate first: behaviors accepted against tested detections deployed, per quarter. Then median age of a live detection, share of the rule set with a named owner, and retirements per quarter. Alert counts and rule totals measure activity rather than throughput, and they reward exactly the wrong behavior.
How many detections should a team maintain?
There is no correct number, and vendors quoting one are counting inventory. The honest ceiling is what you can re-test and retire. A useful exercise: count the rules you could re-run against recorded telemetry tomorrow morning without asking anyone for access. That number is your real rule set.
Does automating security detection engineering require version-controlled detections first?
No. Version control makes the work reviewable and is worth having, but it is not a precondition. The precondition is recorded telemetry you can test a draft against. Without it, automation ships untested logic faster, which is the failure mode the objection to automation is actually about.
When should a detection be retired?
When its telemetry source stops reporting, when the technique it covers has been superseded, when a control now blocks the behavior outright, or when repeated review shows that the rule no longer provides useful or relevant coverage. Retirement is reversible. Keep the logic in the archive, because the technique sometimes comes back.
Stop counting the backlog. Start counting the two numbers that describe the function: behaviors accepted this quarter, and tested detections deployed. If the second is smaller than the first two quarters running, coverage is going backwards no matter how many rules are enabled. The alternative is a conversion that runs continuously, with detections built from new intelligence tested against telemetry you already hold.