SOC

SOC Detection Engineering: Building a Scalable Detection Program

Five criteria for deciding which detection work gets funded, four SOC detection engineering metrics worth reporting to leadership, and how each gets gamed.

Mars Security

Mars Security

Mars Security Research

Forty open detection-gap tickets on the board. One quarter of engineering capacity. Eight of them get built, and the meeting that picks those eight is where SOC detection engineering either becomes a program or stays a queue. This is how that decision gets made, and how you check it afterwards.

SOC detection engineering is the program a security operations team runs to decide which attacker behaviors it will detect, to build and maintain those detections against telemetry it already holds, and to establish afterwards whether the set improved. Three parts. Everybody staffs the middle one, and almost nobody closes the third.

Key takeaways about SOC detection engineering

  • A detection program scales when the criteria that pick the work and the metrics that judge it are the same conversation. Many programs still struggle to close that loop consistently.
  • Capacity is the constraint that makes a SOC detection engineering workflow real. Forty gap tickets against one quarter of build-and-own capacity forces a ranking, and the ranking is the program.
  • Useful detection engineering metrics include time from gap to deployment, the source of newly discovered gaps, and whether deployed detections are being actively reviewed and retired when appropriate. Each needs context because each can be gamed.
  • Every measure of detection effectiveness is biased toward what somebody already thought to look for, so read the detection engineering feedback loop as a trend line, never as a guarantee.

What Is SOC Detection Engineering?

In most SOCs the function arrives as a request queue with a person attached to it. Something appears in the news, a ticket appears. An auditor asks a question, a ticket appears. A program is the standing decision underneath that queue: which parts of your attack surface get watched, taken against a fixed capacity, and revisited when the evidence changes.

The distinction matters because rule sets go static by default. Nobody decides to stop maintaining a detection. Ownership lapses, a log source changes shape, the rule keeps returning nothing, and eighteen months later the set describes an environment you no longer have. A retirement policy and a working feedback loop are the only things that interrupt that drift, because they force someone to look.

How to Build an Effective SOC Detection Engineering Workflow

The part of a SOC detection engineering workflow that carries the program is whatever lets you compare two unrelated pieces of work against each other. Four things do that.

  • A single intake queue. Gaps arrive from hunts, incident retrospectives, audit findings, intelligence reporting, and red team debriefs. When each lands in its own tracker, nothing can be ranked against anything else, and the loudest source wins.
  • A ticket that names the behavior. Every gap ticket states the attacker behavior in plain terms, the telemetry that would show it, and what is missing right now. A ticket that names a product or a report instead of a behavior cannot be scored.
  • A capacity number you publish. How many detections can this team build, tune, and own per quarter, counting the ones already deployed? Without that figure, prioritization is guesswork. With it, every yes names a no.
  • A named owner and a review date on every deployed detection. A person, and a date when somebody has to argue for keeping it. Ownership should resolve to a currently accountable person or team and be reviewed whenever responsibilities change. The program question is narrower: can you see all the pending work in one place, and can you say what you gave up when you picked one item over another.

Prioritizing Which Threats and Detection Gaps to Address

Take the case against prioritizing at all first. Plenty of capable SOCs work purely reactively, building detections off real incidents and real audit findings. That model is evidence-driven, it explains itself to a board in one sentence, and it never spends a quarter on a threat that never showed up. The only argument against it is that it's always exactly one incident late.

That single flaw is why a program needs criteria that survive a triage meeting.

  • Breadth of the behavior. Does it appear across several intrusion paths you care about, or in one campaign? Detections written against a technique in the ATT&CK knowledge base outlive detections written against a specific tool.
  • Telemetry you already hold. If the data does not exist, this is not a detection ticket. It's a logging project, and it belongs in a different queue.
  • Cost to own over cost to build. A detection for bulk export from a SaaS admin console takes an afternoon to write and two years to keep honest, because quarterly compliance pulls and offboarding exports look identical to it.
  • Consequence if it fires late. Identity and control-plane behaviors leave you minutes. Late-stage collection behaviors leave you longer. Rank by the window you get.
  • Overlap with what you already run. A gap that a coarser detection already catches at lower fidelity is a tuning job. Tuning jobs inflate rule count while hiding it. Intelligence reporting gets the same treatment as every other source. A report that produces a ticket nobody ranks is intel that never becomes a detection, and that failure happens in the meeting rather than in the feed.

Two tickets in that meeting look urgent and lose. The first is a critical-severity gap raised after a vendor advisory, for a behavior your environment produces no telemetry for. It scores high on consequence and zero on criterion two, so it becomes a logging request with a date. The second came from an audit finding and duplicates coverage you already hold at a coarser grain. A compliance deadline makes it the loudest ticket in the room. It still loses, and the honest answer to the auditor is the existing detection plus a tuning ticket.

Measuring Detection Effectiveness and Closing the Feedback Loop

The objection to detection engineering metrics is correct, so take it first. Rule count, coverage percentage, and mean time to anything can all be gamed inside two weeks. Write three narrow rules instead of one broad one and the count goes up. Map at tactic level instead of technique level and the percentage goes up. Neither changes what you would catch.

A detection program that only counts what it built is measuring its own effort, not its own effect.

The effect has an outside measure, and it is uncomfortable. Verizon's 2025 Data Breach Investigations Report, in its VERIS discovery method and timeline section, reports that actor disclosure corresponds to 96% of the discovery methods in its dataset. The breach became known because the attacker posted the victim.

Verizon is explicit that the figure reflects where its ransomware data comes from rather than a clean census. Read it as a ceiling on optimism anyway. Being told is the normal way a compromise surfaces, and that is the outcome a detection program exists to change.

MetricWhat it tells youHow it gets gamed
Detections retired per quarterWhether anyone holds the deployed set to a standardRetire quiet, unused rules and leave the noisy debt
Gaps found by hunts versus gaps found by incidentsWhether you are ahead of your own incidentsRun shallow hunts that confirm what you already detect
Days from gap identified to detection deployedWhether the backlog movesOpen the ticket on the day the work starts
Detections that fired at least once in twelve monthsWhether you are carrying dead logicKeep thresholds loose so everything fires occasionally

Where this falls apart. You can't measure what you missed. Every proxy above is biased toward what somebody already thought to look for. Red team results are bounded by the scenarios you commissioned. Hunt findings are bounded by the hypotheses your hunters wrote down. Retrospectives cover the incidents you noticed. Read these numbers as a trend line. They cannot carry a coverage claim.

The loop closes when the measurement output becomes the prioritization input. Retired detections free capacity, so they change next quarter's ranking. Gaps found by hunts change it too, because that is evidence you generated rather than evidence an attacker generated for you. A program that wires those ends together treats detection as a continuous posture instead of a quarterly project. Most never do.

How Mars Security Helps SOC Teams Continuously Improve Detection Coverage

That's what we build at Mars Security. Mars holds a live view of which attacker behaviors your existing telemetry can answer and which it cannot, across SIEM, EDR, identity providers, and cloud sources, and it refreshes as the environment and the intelligence change.

The claim is about cadence. Your coverage picture stops being something you assemble for a quarterly review and becomes something you can read on a Tuesday. The criteria stay yours. What you fund next is still a judgment about your capacity, your consequences, and your risk appetite.

Frequently Asked Questions About SOC Detection Engineering

How does SOC detection engineering differ from the detection engineering practice itself?

The practice is the craft of building and testing a detection. The program is the layer above it: who decides what gets built, against what capacity, and how the result gets judged. A SOC can be excellent at the craft and still run an open loop.

What detection engineering metrics should a SOC report to leadership?

Report a small set that shows both throughput and health: days from gap identified to detection deployed, backlog relative to available engineering capacity, internally surfaced gaps versus incident-discovered gaps, and the share of deployed detections reviewed or validated within their expected window.

How often should a detection engineering feedback loop run?

Monthly for what fired and what stayed silent, quarterly for re-ranking the backlog against capacity, and once a year for a retirement sweep. The annual sweep is the one teams skip, and the one that keeps a rule set from going static.

What detection engineering best practices should carry from 2025 into 2026?

Two hold up regardless of tooling. Fund detection work against a published capacity number, and give every deployed detection a human owner and a review date. Both cost nothing, and both survive whatever your stack looks like next year.

How do you decide a detection is worth retiring?

Ask whether the detection still addresses a relevant threat, whether its telemetry remains healthy, whether the logic still works when tested, and whether its operational cost is justified. A rule that repeatedly produces only low-value alerts may need tuning or redesign; a rule that has never fired should be validated rather than retired solely because the underlying behavior has not been observed.

Who should own the detection backlog in a SOC?

The backlog needs a clearly accountable owner or decision-making group with authority to rank work against available capacity. Detection engineering, threat hunting, and incident response can all contribute candidates, but they should use the same prioritization criteria rather than maintaining competing rankings.

Stop reporting how many detections you shipped this quarter. Start reporting how many you retired, how many gaps your own hunts found before an incident did, and how long each one waited. A program with coverage measured continuously closes that distance faster than one that discovers it in April.