Detection as Code: How to Build, Test, and Validate Security Detections
Learn how detection as code moves rules through Git, code review, and CI, what a detection test actually proves, and why test data is the hard half.

Mars Security
Mars Security Research
A reviewer left one comment on the pull request: which log source is this? Fifteen lines of YAML, a technique everyone in the room could name, and a question that held the rule out of the merge queue for three days. That question is the entire argument for detection as code.
Key takeaways about detection as code
- Detection as code moves rule authoring into a version-controlled repository, so every change to detection logic carries a named author, a readable diff, a recorded test result, and a revert path.
- Version control and CI are the cheap half of the work. The expensive half is test data that resembles production, and every available source of it is incomplete in a different way.
- A rule can compile, pass review and deploy without error while still matching nothing, because the field it depends on is missing from a large share of the fleet.
- Testing before merge proves the logic. Re-running that same test after deployment is the only thing that proves the detection is still alive.
What Is Detection as Code (DaC)?
Detection as code is the practice of managing detection logic the way a team manages application source: rules live in a version-controlled repository, changes arrive as pull requests, automated checks run before merge, and deployment into the SIEM or EDR happens from the pipeline rather than from a console.
Three things follow, and together they are the benefit. Every rule has an author you can name. Every change has a diff you can read. Every deployment has a commit you can revert. A team with all three can answer when this rule changed and who approved it without asking anyone.
Detection as code isn't a tooling choice. It's a change-control model for logic that runs against production data, borrowed from software engineering, which solved it while security was still editing live systems by hand.
How Modern Detection Engineering Works
The workflow runs in six steps. The interesting failures all happen between four and six.
- A behavior gets chosen from intelligence, from a hunt result, or from an incident.
- Someone writes the rule against a named log source.
- A pull request opens against the detection repository.
- CI compiles the rule and replays recorded events against it.
- A human reviews the logic, then separately reviews the coverage claim.
- Merge deploys the rule, and validation continues on a schedule.
Step one is a different discipline, and this article starts after it. If your problem sits earlier, the harder question is turning intel into a detection rather than how to store one.
The rest is not theoretical. You can count it yourself. On September 7, 2026, the public SigmaHQ rule repository carried 16,885 commits on master, and elastic/detection-rules carried 4,066 on main. Two large rule sets, both maintained in public through pull requests.
One rule runs through the rest of this article: a Sigma rule for credential access against LSASS via rundll32.exe and comsvcs.dll, which MITRE tracks as T1003.001. Fifteen lines. It fails CI.
Traditional vs. Code-Based Detection Workflows
The console workflow deserves a fair hearing, because it wins on the axis that matters during an incident. An analyst who can see live data writes a query, eyeballs the hits, tightens the noisy clause, and saves it. Twenty minutes, start to finish. No branch, no reviewer in another time zone, no pipeline run. For a four-person team working an active intrusion, that speed beats reviewability, and anyone who says otherwise has not worked that Sunday.
| Console workflow | Detection as code | |
|---|---|---|
| Time to first deployed rule | Minutes | Hours to days |
| Change history | A last-modified timestamp | A diff, an author, a review |
| Review | Optional and informal | Enforced by branch protection |
| Test before deploy | The analyst's own eyes | Fixtures replayed in CI |
| Rollback | Rewrite it from memory | git revert |
| Multiple backends | Rewrite per platform | Portable formats can compile to multiple backends |
| Characteristic failure | Undocumented drift | Ceremony without validation |
Detection as code buys governance and pays in latency. A team that pretends the latency is free goes back to the console during the first bad week.
Neither column addresses decay. A rule pinned to a hash or a file name degrades on the same schedule in a console or a repository, which is why static rules degrade wherever you keep them.
Building a Scalable Detection Workflow
Four directories carry most of the weight. Their names matter less than each having one job.
- detections/ - one rule per file, grouped by log source rather than by threat actor, because log source is what breaks.
- tests/fixtures/ - the events each rule must match, and the events it must not.
- pipelines/ - the field mappings that translate a rule into each backend's schema.
- CODEOWNERS - the person or team responsible for reviewing changes to that detection or log source.
Branch protection is the part teams skip and the part that makes the rest real. One approving review, CI green, no direct pushes to main. No console edits. No undocumented changes. No rule whose author left in 2024.
Now the common objection: teams try this, find the pipeline costs more than it returns, and watch rules break in production anyway. Right about the symptom, wrong about the cause. Abandoned programs almost always built the gates first - schema validators, naming conventions, a template with eleven required fields. Tidy repository, untested rule set.
Spend the first sprint on fixtures instead. Linting can be added in an afternoon a year from now. Fixtures cannot: six months on, nobody remembers what the rule was supposed to match, and the events that would have proved it have aged out of retention.
Using Git, CI/CD, and Automation for Detection Engineering
The pull request adds one file, at detections/windows/process_creation/lsass_dump_comsvcs.yml:
title: LSASS credential access via comsvcs.dll MiniDump
status: experimental
logsource:
category: process_creation
product: windows
detection:
selection:
Image|endswith: '\rundll32.exe'
CommandLine|contains|all:
- 'comsvcs.dll'
- 'MiniDump'
condition: selection
falsepositives:
- Administrative crash-dump collection
level: high
The job that runs against it does two things, and the second is the one that earns its place:
jobs:
detections:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: pip install sigma-cli pytest
- run: sigma plugin install splunk sysmon
- name: Compile every rule
run: sigma convert -t splunk -p sysmon detections/ -o /tmp/rules.spl
- name: Replay fixtures
run: pytest tests/ -q
The first step proves the rule can be parsed and converted for the target backend using the configured pipeline. It catches syntax, conversion, and some mapping problems, but it does not prove that the required fields are populated in production. The second replays recorded events and asserts what matched.
The reviewer's comment arrived before the job finished. CommandLine is populated on Sysmon Event ID 1. On Windows Security 4688, it is populated only where command-line auditing has been switched on by policy. The rule does not say which of those it assumes, and the fleet contains both.
A detection in version control is auditable. It is not thereby correct.
How to Test and Validate Detection Rules
Four layers, in ascending order of cost and descending order of how often teams build them.
- Schema and syntax. The rule parses, the required fields are present, the modifiers exist. Free, fast, catches almost nothing interesting.
- Compilation. The rule converts to the target query language using the configured field mappings. This catches conversion and mapping errors, but does not prove that the mapped fields exist or are populated across your production environment.
- True-positive fixtures. Recorded events the rule must match, one per delivery path you claim to cover.
- True-negative fixtures. Recorded benign events the rule must not match, drawn from the noisiest hosts you own.
Here is what CI returned:
FAILED tests/test_lsass_dump.py::test_matches_4688_event
rule: detections/windows/process_creation/lsass_dump_comsvcs.yml
fixture: tests/fixtures/4688_rundll32_comsvcs.json
expected 1 match, got 0
reason: field 'CommandLine' not present in event
The rule is fine. The fixture is fine. What failed is an assumption the author never wrote down, and the test is the first artifact capable of catching it. No expression recovers a field the agent never sent, so what the reviewer asked for is a declaration:
logsource:
category: process_creation
product: windows
+ definition: >
+ Requires the process command line in the event. Sysmon Event ID 1,
+ or Windows Security 4688 with command-line auditing enabled.
That converts a silent coverage gap into a documented dependency. Somebody now decides: enable command-line auditing across the fleet, or accept that this detection covers part of it.
Here is where testing gets uncomfortable. The hard part is usually not the harness itself; it is building and maintaining representative test data. It is the test data, and every source of it is compromised differently.
| Test data | Best for | What it will not tell you |
|---|---|---|
| Hand-written synthetic events | Field logic, modifiers, edge cases, negative tests | Whether the field exists in your fleet at all. Synthetic events pass rules that production breaks |
| Replayed production events | Field presence, real noise, genuine benign look-alikes | Anything about a technique not yet used against you. It carries secrets, and ages out of schema fast |
| Adversary emulation, such as Atomic Red Team | End-to-end proof on real hosts, execution through to alert | Anything nobody has written an atomic for. It covers the techniques someone already thought of |
Use all three and hold the split honestly: synthetic for logic, redacted production for field presence, emulation for end-to-end proof. No team has this fully solved. The ones furthest along keep a redacted production sample with a named owner, and they will tell you it is the least popular ticket on the board.
Moving From Automated Testing to Continuous Validation
A merged rule works as of one commit. That is the whole guarantee, and it expires quietly.
The agent upgrades and renames a field. A log source gets dropped for cost. An exclusion added during a noisy week is never removed. None of these produce an error anywhere. They produce silence, and silence reads exactly like a quiet environment.
- Re-run the emulation. Execute the same atomic against a canary host on a schedule, then assert the detection fired inside a set window. This is the CI test pointed at production instead of at a fixture.
- Watch the rule's own hit rate. Zero matches over ninety days is a question, and it deserves a ticket.
- Diff the schema. Compare the fields your rules reference against the fields your log sources currently emit, after every agent upgrade.
Running that loop by hand is a standing tax on a small team, which is why our platform runs it continuously rather than quarterly. The mechanism is identical either way: execute the behavior, look for the detection, treat silence as a finding.
Putting Detection as Code to the Test Against Real-World Attacks
The behavior. To use credentials nobody typed on the machine, an attacker has to read them out of memory. On Windows, that means opening a handle to lsass.exe with rights permitting a memory read, then writing the contents somewhere retrievable. MITRE tracks it as T1003.001. Routing that through rundll32.exe and comsvcs.dll is one delivery path among several, popular because both binaries are signed and already present.
The logs. Sysmon Event ID 1 gives you Image, CommandLine and ParentImage. Sysmon Event ID 10 gives you the process-access view: which process opened lsass.exe, and with what rights. Windows Security 4688 gives you the execution event, with the command line only where auditing is on. The rule reads the first source. The hunt reads the second.
The hunt. List every process that opened a handle to lsass.exe with memory-read rights over thirty days, grouped by opening image and counted by host. Most of the output is a short, stable list of expected openers. What you want is the tail: one host, one signed binary that has never opened lsass.exe before, once.
Coverage. An attacker who calls the dump routine from a compiled binary never touches rundll32.exe. One who copies and renames the binary defeats the Image|endswith clause. One working on a 4688 host without command-line auditing produces an event the rule cannot read at all. The process-access hunt catches the first two and is blind to the third.
The benign look-alikes are why this needs a hunt rather than an alert. Windows Error Reporting collects crash dumps. Endpoint protection, DLP, and memory-scanning agents read lsass.exe constantly and legitimately. IT captures a dump when a host hangs. On any real fleet, the expected-opener list is longer than the detection, and building it is the work.
Which returns to the comment that opened this. The repository gave that rule an author, a review, a test result and a revert path. It did not make it durable. A brittle rule under version control is a brittle rule with better paperwork, and the paperwork earns its keep only by making the next question askable: what would have to change here for this rule to stop firing, and would anyone notice?
Frequently Asked Questions About Detection As Code
What tools do teams use for detection as code?
Git hosting plus a CI runner is the whole requirement; the rest is preference. Most teams pair a rule format, a converter, a test runner in a language they already write, and the platform's deployment API. Several SIEM and EDR vendors now ship first-party pipelines of their own.
How is detection as code different from detection engineering?
Detection engineering is the discipline: deciding what to detect, from which intelligence, at what fidelity. Detection as code is one delivery model for it. You can practice detection engineering badly with a console and a spreadsheet, and you can run a spotless repository full of detections nobody chose deliberately.
Do I need Sigma to do detection as code?
No. Sigma pays off when you deploy to more than one backend, because one rule compiles into several query languages. If you run a single SIEM and expect to keep running it, writing rules in that platform's native format and testing them in CI delivers most of the benefit with less translation risk.
How many rules justify building a pipeline?
There is no clean threshold, but two signals are reliable: the first time two people edit the same rule in a month, and the first time nobody can explain why an exclusion exists.
Who should own the detection repository?
The team that gets paged. Ownership drifts toward whoever built the tooling, which puts the repository with a platform engineer who never sees the alerts and has no reason to care about fidelity. Name a person per log source in CODEOWNERS, and revisit that file whenever people change roles.
Does this apply to EDR and cloud detections, or only SIEM rules?
Anywhere the platform exposes its rules through an API, the same treatment applies: custom EDR detections, cloud provider rules, identity policies. The limit is vendor-managed logic you cannot export. Built-in EDR detections stay a black box, and your repository can only record which ones you enabled and when.
Stop treating a merged pull request as a shipped detection. Start treating any rule that has not fired since the day it merged as an open question. The repository, the review and the CI job are the parts you can finish inside a quarter; validating detections continuously is the part with no end state, and the only part that tells you whether the rest of it worked.