Detection Engineering

SIEM Detection Engineering: A Guide to Building Better Detections

Learn why SIEM detection rules fail in production, what suppression really does in Splunk, Microsoft Sentinel, and Elastic, and how to tune without blind spots.

Mars Security

Mars Security

Mars Security Research

The rule passed review on a Thursday. By Monday it owned the alert queue, and someone had quietly written a filter that sent its output to a folder nobody reads. Most SIEM detection engineering failures look like that one, and the query is rarely the reason.

It happened because the data the rule depends on behaves differently in production than it did in test. One field arrives in two formats from two collectors. One host writes local time. One batch job performs the attacker's behavior every night at 02:00 and has done so for three years. The logic held. Your model of the environment did not.

Key takeaways about SIEM detection engineering

  • Many production detection failures come from data problems: fields that are absent, inconsistent, delayed, or populated differently across sources.
  • Noise is often a correct rule meeting an environment nobody documented. Find the legitimate process producing the behavior before you rewrite the query.
  • Suppression is not one behavior. Splunk, Microsoft Sentinel, and Elastic each stop something different when a rule is throttled, and each opens a different gap.
  • Record what every exclusion and suppression window costs in coverage, with the date you will revisit it. Published work suggests exclusions often persist for years after they are added.
  • Measure what your SIEM already reports, and nobody opens: rule execution gaps, failed runs, and how many rules produced a true positive last quarter.

1. Why SIEM Detections Fail in Production

A reviewed rule fails after deployment for a small number of reasons, and they repeat across every platform:

  • Field divergence. Two sources populate the same field with different values, formats, or truncation, and the rule was written against one of them.
  • Clock disagreement. The host writes local time, the index stores UTC, and the correlation window that made sense in test drifts by an hour twice a year.
  • An undocumented environment. A batch job, a migration, a contractor's deployment tooling, or a scanner does exactly what the rule was written to catch.
  • Late data. The event arrives after the look-back window has already closed over the period it belongs to.
  • A rule that stopped running. Disabled during an incident, timed out under load, skipped when the scheduler saturated. Nothing alerts on nothing. Only the first two are query defects. The third is especially important because it can generate large volumes of legitimate matches and often attracts the wrong fix.

The rule described real behavior. The environment contains a legitimate producer of that behavior that nobody wrote down. That isn't a logic problem. It's a modeling problem, and treating the first week of noise as a discovery exercise rather than a defect ticket gets you a better rule and a better inventory, which is a large part of why rules degrade in production even when nobody touches them.

One constraint sits underneath all five. A rule sees only what was ingested, normalized, and indexed, and each step drops something: what the collector was not configured for, what the schema has no field for, and what arrived after the window closed. None of it is visible from the search bar.

2. Choosing the Right Security Data for Each Detection

Data selection wins or loses fidelity before the first line of logic. Pick the source that carries the behavior in a form consistent across every host, account, and hour of the day, and run each candidate field through three checks: is it populated everywhere, is it populated identically, and does it arrive in time.

Data sourceThe field the rule leans onWhat to verify before you trust it
Windows process creationCommandLineAudit policy is set per host and drifts. Confirm population across every host group, not a sample
EDR process telemetryparent_processSensors reconstruct lineage differently. A re-parented or injected process breaks the join
Identity provider sign-in logsthe user principalUPN in one source, SAM account name in another, an object id in a third
Cloud audit logsthe caller identityA call one service makes for another does not carry the human who started it
Proxy and network logssrc_ipOften the load balancer, unless the forwarded header is parsed into its own field

Delivery latency needs its own check because it fails silently. Microsoft's guidance on ingestion delay in Sentinel states the consequence plainly: if an event is generated inside the look-back period but ingested after that run, the next run's time filter drops it as too old, and the rule does not fire an alert (Microsoft Learn, Handle ingestion delay in scheduled analytics rules, updated August 2026). The documented fix extends the look-back by the measured delay, then re-anchors the query with ingestion_time() > ago(rule_look_back) so the extension does not duplicate alerts.

3. Building High-Fidelity Detection Logic Across SIEM Data

Take one rule and follow it: a new process spawned under a service account. Good detection, validates in a lab in ten minutes, drowns most teams who deploy it as written. Production logic carries three things the lab version did not.

Anchor on identity, not on host. Service accounts move. A rule scoped to a host group stops covering the account the moment the workload is rescheduled onto a new node.

Make the join key explicit. If the account arrives as a UPN from the identity provider and a SAM account name from the endpoint, normalize one to the other inside the query.

Own the time base. Extend the look-back by the measured delay, then use the ingestion-time restriction so the extension does not fire the same event twice.

let lookback = 10m;

let delay = 5m;

let ServiceAccounts = datatable(Account:string)["svc-backup", "svc-deploy"];

let ApprovedParents = datatable(Process:string)["ccmexec.exe", "MonitoringHost.exe"];

DeviceProcessEvents

| where TimeGenerated >= ago(lookback + delay)

| where ingestion_time() > ago(lookback)

| where AccountName in (ServiceAccounts)

| where InitiatingProcessFileName !in (ApprovedParents)

| project TimeGenerated, DeviceName, AccountName,

FileName, ProcessCommandLine, InitiatingProcessFileName

Then name what legitimately looks identical: patch deployment under a service identity, backup agents spawning helper binaries on a schedule, authenticated vulnerability scans, configuration management runs, and the monitoring agent restarting a failed child several hundred times a day. Every one is a real producer of new-process-under-service-account, and every one belongs in the rule's documentation before it reaches the queue.

4. Testing Detections Against Real Attacker Behavior

A test that fires the rule proves it is syntactically alive. It proves nothing about production. Three tests do.

Replay history. Run the finished logic across thirty to ninety days of stored data and count the results. That number is your first-week alert volume, available before deployment instead of after. Four figures means the rule is not ready.

Test the negatives on purpose. Point the rule at the windows when the backup runs, the patch cycle fires, and the scanner sweeps. A rule that cannot separate those from the behavior you care about needs a different data source, not a tighter filter.

Change one attacker-controlled string. Rename the binary, move the path, swap the parent. If the rule stops firing it was matching an artifact, which is the failure mode that matters most against attacker behavior that adapts faster than a rule set can be edited.

Record all three results in the rule itself. The historical count, the named benign producers, and the surviving variants are the only defensible basis for the tuning decisions that follow.

5. Tuning Detection Rules Without Creating Blind Spots

Every tuning decision is a coverage decision. Most teams only write down the first half.

You add an exclusion. The blind spot is everything else that satisfies it. A path exclusion covers whatever can be written to that path, and it outlives the process that justified it. A preprint posted to arXiv on 31 August 2026, The Exclusion Ratchet, measured this across nine years and 8,234 revisions of the SigmaHQ corpus and estimated that 86.7% of exclusions remain in force three years after they are added, with persistence unrelated to whether the rule was the only coverage for its technique.

You raise the threshold. The blind spot is the patient version of the same behavior. Ten failures in a minute alerts; three an hour for a week does not.

You suppress or throttle. The blind spot depends entirely on the platform, and the three most common SIEMs do three different things:

  • Splunk throttling suppresses alert triggering for a set period, and per-result throttling suppresses on the fields you name. Suppression groups go further: when one alert in the group fires, the whole group is suppressed. Editing the alert, its schedule, or the throttle configuration removes the throttle file, so a throttle you believe is running may not be (Splunk Enterprise 10.4 Alerting Manual, updated May 2026).
  • Microsoft Sentinel offers "Stop running query after alert is generated" for up to 24 hours. That setting stops the query, not the alert. Whatever happens inside the window is never evaluated (Microsoft Learn, Create a scheduled analytics rule from scratch, updated July 2026).
  • Elastic groups matching events into one alert and keeps the count in kibana.alert.suppression.docs_count, so volume survives. Suppressed alerts still count toward the rule's alert ceiling, which for threshold, event correlation, ES|QL and machine learning rules is the Max alerts per run setting, default 100; custom query rules have no suppression limit. Suppression requires the right subscription tier (Elastic Docs, Suppress detection alerts). You drop a log source. The blind spot is everything only that source saw, including the detections you have not written yet.

Where this falls apart: there is no tuning without loss. Suppressing a noisy source removes coverage of whatever else that source would have shown, and no framework recovers it. The only defensible practice is to write down what each suppression costs and revisit it on a schedule. Six columns do the job: rule, field, value, who approved it, what coverage it removes, and the date it gets re-examined. Most teams keep no such register, and the ones that start usually find something uncomfortable inside a quarter.

6. Measuring and Improving SIEM Detection Coverage

Coverage usually gets reported as a count of rules mapped to techniques, which measures the rule set rather than the detection. Four measurements your SIEM already produces sit closer to the truth.

Execution gaps. Elastic reports a gap duration per rule: how much source event time was not searched between one run and the previous one, with documented causes including the rule being disabled and the rule failing to execute at all. Coverage loss, expressed in minutes.

Failed runs. A rule that times out under load has zero coverage for that interval and reports no alerts, which looks identical to a quiet night.

The suppression register. Every open suppression is a documented hole. Count them, age them, and read the ones older than the current threat model.

True-positive contribution. Track which rules have produced confirmed useful detections over time, but interpret zero-hit rules alongside threat relevance, testing results, and telemetry health rather than treating silence as failure. It is one useful prioritization input alongside ATT&CK mapping, validation results, and threat relevance.

For MITRE ATT&CK SIEM use cases, map the rules that are running and unsuppressed, not the rules that are installed. A technique mapped to one heavily excluded or frequently suppressed rule may have substantially less effective coverage than the matrix suggests.

7. Building a Continuous SIEM Detection Engineering Workflow

The loop is short enough to state without a diagram. A working SIEM detection engineering workflow runs seven steps, and the arrows all return to the start:

  • Intake. A behavior worth detecting arrives from intelligence, an incident, or a hunt.
  • Data check. Confirm the fields exist, agree across sources, and arrive inside the schedule. Kill the candidate here if they do not.
  • Author. Write the logic with the join key and the time base made explicit.
  • Replay. Run it across stored history, count the results, name the benign producers.
  • Deploy with an expiry. Every rule enters production carrying a review date.
  • Measure. Gaps, failures, suppressions, and true-positive contribution, on a fixed cadence.
  • Retire or rewrite. A rule producing nothing but noise for two quarters is not a detection. New techniques enter at step one and travel the same path as everything else. The teams that keep up are not the ones writing rules fastest. They are the ones whose step two rejects candidates their telemetry cannot support, and whose step six runs whether anyone remembers to schedule it or not.

Frequently Asked Questions About SIEM Detection Engineering

Why do SIEM detection rules generate false positives after they pass testing? Because test data is clean and production data is not. Fields that were consistent in the lab arrive from several collectors in different formats, timestamps disagree, and legitimate automation performs the same behavior the rule was written to catch. The logic is usually correct; the assumption about the environment is not.

What should a SIEM detection engineering workflow diagram actually show? It should show a loop that closes. The useful version has intake, a data-quality gate that can reject a candidate outright, authoring, historical replay, deployment with a review date, measurement, and retirement feeding back into intake. Most published diagrams omit the gate and the retirement step.

How is suppression different from an exclusion? An exclusion changes the rule's logic permanently, removing a condition from what it will ever match. Suppression is a time-based control on alert output that leaves the logic intact. Both remove coverage, but only the exclusion changes what the rule can see, so the two need separate records.

How often should production detection rules be reviewed? Rules carrying broad exclusions or long-lived suppression should be reviewed more frequently than stable detections, with the cadence based on threat relevance, change rate, and the breadth of the visibility tradeoff.

What belongs in a pre-deployment checklist? Field population across every host group, format consistency between sources, measured ingestion delay per table, a historical replay count, the named benign producers, the join key you chose, and the review date. Seven items, and any one left blank predicts the failure you will get.

What is the fastest way to reduce noise without losing coverage? Change the data before you change the logic. Adding a field that separates the benign producer from the attacker, such as a parent process or an authentication method, keeps the detection intact. Filtering the benign producer out hands the attacker somewhere to hide.

Stop tuning in the alert console. Start tuning in a record that states what the rule stopped seeing and when you agreed to that, because the register is the only artifact that outlives the engineer who wrote it. That assumption sits underneath our own approach to validating detections continuously rather than once a quarter, and the register costs nothing to start today.