Detection Engineering Lifecycle: From Threat to Detection
The detection engineering lifecycle explained in five stages, with the handoff between each stage named, its owner, and the point where the work usually stalls.

Mars Security
Mars Security Research
A rule has been firing eleven times a day since March. Nobody remembers who wrote it, which report it came from, or what a responder is supposed to do about it, and nobody has turned it off. The distance between a rule existing and a rule earning its place is what the detection engineering lifecycle is for.
Key takeaways about the detection engineering lifecycle
- The loop runs in five stages: prioritize, specify, build, deploy, and measure. The fifth is the easiest to skip, and skipping it flattens the loop into a straight line.
- Every boundary between stages is a handoff with a named artifact and a named owner. The argument here is that work stalls at the boundaries rather than inside the stages, and that stage-by-stage process documents look straight past them.
- Measure the cycle time from a threat being prioritized to its detection running in production. That single number says more about your detection engineering process than a coverage percentage does.
- Detection effectiveness is hard to measure honestly, because it requires knowing what you missed. Hunt findings, red team results, and incident retrospectives are proxies, and each carries a bias worth naming out loud.
What Is the Detection Engineering Lifecycle?
The detection engineering lifecycle is the repeating sequence of work that carries a threat into a running detection and carries production evidence back into the next prioritization decision. A practical way to structure that lifecycle is in five stages:
- Prioritize. Decide which threat behavior gets engineering time, and why this one now.
- Specify. Convert that threat into a detection requirement: the observable behavior, the telemetry source it lands in, and the benign activity that will look identical.
- Build. Write the logic, then test it twice. Once against the attacker behavior, once against ordinary traffic.
- Deploy. Move it into production with an owner, a severity, and a response a tired analyst can follow at 2am.
- Measure. Judge what it caught, what it missed, and what it cost, then hand that back to stage one. Many detection engineering lifecycle models include a final stage for monitoring, tuning, validation, or continuous improvement. Naming it is not running it. The fifth is what makes this a lifecycle: without a return path, rules accumulate, none retire, and the rule set turns sedimentary.
This isn't a diagram. It's a loop with a cycle time you can measure, from the day a behavior gets ranked to the day its detection runs in production.
The case for not doing any of this. For a very small team, a heavily formalized lifecycle can create more process than value. A ticket queue and shared repository may be enough while ownership and rationale remain obvious to everyone involved. That argument stops holding once ownership, history, or rationale can no longer be reconstructed reliably.
Prioritizing Threats and Identifying Detection Requirements
Stage one answers a single question: of everything that could hurt you, which behavior gets engineering time this week? The inputs are ones you already hold. An intel report. An incident you closed last month. A gap someone tripped over during a hunt. None of them arrive ranked, and ranking them honestly is the step a busy week quietly drops.
When it does get dropped, the item picked is the one with the cleanest log source and the shortest query, which is a scheduling decision wearing a risk decision's clothes. Rank on what an attacker has to do inside your environment, the same reason why detection outlasts remediation.
The input side also grows on a schedule you do not control. Enterprise ATT&CK carried 216 techniques and 475 sub-techniques when v18 published on October 28, 2025. Six months later, v19 published on April 28, 2026 with 222 techniques and the same 475 sub-techniques, and it split the Defense Evasion tactic into Stealth and Defense Impairment, relabeling mappings teams had already made. In August 2026, MITRE introduced its first Agile release, a narrower-scope model for publishing targeted updates to Groups, Software, and Campaigns between the standard biannual releases when significant threat activity warrants it.
Stage two hands stage three a requirement with four fields, never a ticket reading "detect lateral movement":
- **The **behavior, written as something the attacker has to do, not as a tool name.
- The telemetry, named down to the source and the field, and confirmed to exist before anyone writes a line of logic.
- The benign look-alikes, listed at requirement time rather than discovered in production.
- The decision, meaning what a responder does when this fires. Take one unglamorous example the rest of the way: a service account authenticating from an ASN it has never used. Behavior, credential use from new infrastructure. Telemetry, the identity provider's sign-in events, with autonomous_system_number and service_principal_id. One of those fields has to be enriched before anything can query it, and finding that out now rather than in stage three is the entire reason stage two exists.
Designing, Building, and Testing Detection Logic
Stage three is where the detection engineering workflow earns its reputation for being software work. Version control, review, and continuous integration sit underneath everything below. The part specific to detection is that the logic gets tested twice, against two different populations.
- Against the behavior. Reproduce the authentication from new infrastructure in a controlled account and confirm the logic fires.
- Against a normal week. Run it backwards over production history and count what it would have produced. This is easy to skip, but it is critical for estimating how much benign activity the detection will produce in production. For the service account example, the second test is unkind. A cloud provider moves the workload between regions and the egress ASN changes. A vendor rotates its IP ranges. The network team finishes a migration on a Sunday and every service account in the estate looks new on Monday. All three are legitimate, and all three are indistinguishable from the thing you are hunting if the ASN is the only field you look at. That is why the requirement pairs it with the identity's own history rather than a global allowlist.
The envelope handed to stage four contains the logic and the test cases. Both. Nobody except its author can safely change a detection that shipped without the benign cases used to tune it, and its author will eventually leave.
Deploying Detections and Measuring Their Effectiveness
Stage four is easy to treat as the finish line. A detection goes live with three things attached: an owner who answers for it, a severity that matches what it means, and a written response. Detections handed to the SOC without the third become noise, and noise is how a rule set earns the right to be ignored.
Stage five decides whether any of the previous four were worth doing. It is also the hardest of the five to do honestly.
An honest caveat. Measuring detection effectiveness requires knowing what you missed, and by definition you do not. Every substitute is partial. Hunt findings that no rule caught lean toward whatever the hunter chose to look for that quarter. Red team results lean toward the red team's tradecraft and their scope document, and they are scheduled, which real adversaries are not. Incident retrospectives lean toward incidents somebody noticed, a survivorship problem sitting at the center of your evidence.
Use all three, name the bias each time you cite one, and treat the aggregate as directional.
If four stages are operating but the loop still does not close, measurement and the return path into prioritization are the first places to inspect: nothing breaks visibly when it stops. Without it, stage one keeps guessing. Closing that return path is what turns detection work from a quarterly project into detection as a continuous posture.
| Stage | Input | Output handed on | Owner | Where it stalls |
|---|---|---|---|---|
| 1. Prioritize | Intel, incidents, hunt gaps, risk register | A ranked backlog item naming the threat behavior and why now | Detection lead | Ranked by what is easy to write instead of what is likely |
| 2. Specify | The ranked backlog item | Behavior, telemetry source and fields, benign look-alikes, response decision | Detection engineer, with intel input | Written before anyone checks the log source exists |
| 3. Build | The requirement | Tested logic plus the test cases, true positive and benign | Detection engineer | Tested against the attack only, never against a normal week |
| 4. Deploy | Logic and test cases | A running detection with owner, severity, and response | Detection engineer to SOC | Shipped with no response guidance, so it becomes noise |
| 5. Measure | Production outcomes, hunt findings, red team, retros | Evidence that changes the ranking in stage one | Detection lead with SOC lead | No route back to stage one, so the loop is an arrow |
How Mars Security Creates a Continuous Detection Engineering Lifecycle
The five stages do not change when the tooling changes. Cadence does. A lifecycle that runs once a quarter produces a rule set reflecting last quarter's threats, and that gap is the whole problem.
That is what we build at Mars Security: the same five stages, running continuously rather than in project form. Intelligence converted into detections as it arrives, coverage mapped against what is already deployed, and hunts running against your telemetry where it already sits. No ingestion pipeline. No rip-and-replace. No new headcount.
Frequently Asked Questions About the Detection Engineering Lifecycle
How do I measure the cycle time of my detection engineering lifecycle?
Timestamp two events: the date a behavior enters the ranked backlog, and the date its detection starts returning production evidence. Track the median across recent items to understand typical cycle time, but also watch a high percentile such as the 90th or 95th to expose items stuck at handoffs. Track the trend over time rather than optimizing for one absolute number. Track the trend, not the absolute number.
What is a detection engineering maturity model, and do I need one?
A detection engineering maturity model scores a program across dimensions like coverage, testing rigor, automation, and feedback, then places it on a ladder. It is useful for arguing a budget case. It is a different instrument from a lifecycle, which describes the work itself, and you can run the loop competently without ever scoring yourself.
Who owns the lifecycle when there is no dedicated detection engineer?
Someone still owns each handoff, whether or not the title exists. In a small team, one person may own several stages. The important requirement is that every handoff has an explicit owner, even when the same analyst appears on both sides of it. What breaks a program is a stage with no named owner at all, because unowned stages are where items sit indefinitely.
What makes a good detection engineering pipeline?
Judge a detection engineering pipeline on what it hands between stages rather than on how many rules it produces. Requirements name their telemetry before logic gets written, detections ship with their test cases attached, and production evidence has a defined route back into prioritization. Throughput without that return path only accumulates faster.
How is this different from the threat intelligence lifecycle?
The threat intelligence lifecycle runs from direction through collection, processing, analysis, and dissemination, then feedback on whether the finished product answered the question it was set. The detection engineering lifecycle starts near its dissemination point and ends with running logic plus evidence about how it performed. They overlap at one seam, where a finished intelligence product becomes a ranked detection requirement, and that seam is where handoffs get dropped.
What should happen to a detection that stops firing entirely?
Silence is ambiguous. It can mean the behavior stopped, the logic broke, or the log source quietly went away, and those demand different responses. Check the telemetry is still arriving before you touch the logic. A detection silent through a confirmed source outage was never protecting anything.
Stop treating deployment as the end of the process. Start measuring the interval between a threat being ranked and its detection returning evidence, then run the loop again on what that evidence says, because a lifecycle that runs continuously is the only version of this that ever gets shorter.