← ./resources / blog

Advanced EDR Architecture and Tiers of Detection

Endpoint Detection and Response is not one alert engine. It is a system of sensors, local enforcement, evidence pipelines, analytics, analyst workflows, and containment controls. Mature coverage uses multiple detection tiers so a single missed signature does not decide the outcome.

What an EDR is designed to do

An EDR agent observes endpoint activity, enriches it with context, sends selected telemetry to an analysis service, and gives defenders the ability to investigate and respond. Antivirus-style prevention remains valuable, but EDR extends the job: it builds timelines, correlates actions across hosts and identities, preserves evidence, and can isolate an endpoint when risk is high. The goal is not to record every byte forever. It is to capture sufficiently trustworthy, useful events to answer what happened and to stop it safely.

High-level EDR architectureEndpointprocess + image loadsfile, registry, networkscript + identity eventslocal prevention / bufferDetection platformingestion + normalizationcorrelation + scoringintel + behavior modelscase timeline + evidenceSecurity operationsanalyst investigationSIEM / identity contextcontainment decisionisolate / collect / remediateauthorized response action returns to the endpoint
Figure 1. A resilient EDR pipeline turns endpoint facts into decisions, then sends a constrained response action back through the control plane.

Tier 0: prevention and policy enforcement

Tier 0 is the work that prevents known-bad or clearly disallowed behavior before it becomes an investigation: reputation checks, signed-code policy, attack surface reduction rules, script restrictions, exploit protections, and application control. It has the lowest analyst cost and should be measured by blocked activity and business-compatible coverage. Prevention needs an exception process; uncontrolled exclusions create blind spots that later analytics cannot reliably repair.

Tier 1: atomic, high-confidence detections

Tier 1 alerts are close to an observable: a known malicious hash, a revoked certificate, an untrusted driver load, or a strong rule with a narrow false-positive profile. They are fast to triage and ideal for automated enrichment. Their weakness is brittleness: attackers can change file names, infrastructure, or tooling. A mature program keeps these detections, but does not confuse them with comprehensive behavior coverage.

Tier 2: behavioral and sequence analytics

Tier 2 looks at relationships and time. Examples include an office application spawning an unusual interpreter, a newly created process immediately making rare external connections, or a service installation followed by unexpected credential access. These detections should describe the behavior in plain language, list required telemetry, and have a response hypothesis. They are more durable than simple indicators, but demand tuning against the environment's legitimate automation.

Tier 3: cross-domain correlation and human-led hunting

Tier 3 combines endpoint, identity, network, cloud, email, and threat-intelligence signals. One event may be benign; a chain across multiple controls can become decisive. Analysts use this layer to ask scoped questions: which hosts executed a newly seen signed binary, what identities authenticated from those hosts, and did any use a privileged path? Hunting converts the answer into a tested detection, a hardening change, or a documented negative result.

Detection tiers: broaden context as certainty requiresTier 0 — prevention and policy enforcementTier 1 — atomic, high-confidence signalsTier 2 — behavioral sequencesTier 3 — correlation + huntingHigher tiers add context and analyst effort; lower tiers apply quickly and broadly.
Figure 2. Tiers are complementary. Reliable defense applies low-cost prevention broadly, then adds behavior and context where it matters.

Telemetry engineering: collect for decisions

High-quality process, command-line, image-load, network, script, registry, service, scheduled-task, and authentication telemetry makes detections explainable. Excess collection without retention planning creates cost and noise; too little makes investigations speculative. Define the questions the team must answer, map each to fields and data sources, test collection on representative endpoints, and monitor agent health as carefully as alert volume.

Safe lab setup with GitHub tools

  1. Use a non-production Windows VM, an approved test account, and a snapshot.
  2. Deploy Sysmon with a reviewed configuration such as sysmon-modular; adapt it to your environment rather than copying rules blindly.
  3. Send events to a lab collector and use Sigma as a portable detection-rule language. Validate each rule against normal business activity.
  4. For authorized coverage testing, use benign simulations and document expected event IDs, alert, analyst decision, and cleanup result.

Response is a safety-critical feature

Containment actions have real business impact. Endpoint isolation, process termination, file quarantine, remote evidence collection, and account disablement should be role-controlled, logged, reversible where possible, and paired with a clear escalation path. An automated isolation decision may be appropriate for a high-confidence ransomware signal; a weak anomaly should create a case, not a disruption. Test response actions during exercises before relying on them in an incident.

Relevant CVEs: visibility must survive product risk

CVE-2021-44228 (Log4Shell) demonstrated how quickly widespread exploitation can turn endpoint telemetry into an essential scoping tool: defenders needed to find affected processes, outbound connections, and follow-on activity. CVE-2024-21412 was a Windows Internet Shortcut Files security-feature bypass used in the wild, showing that initial access controls can fail and endpoints still need behavioral detection. These CVEs are not EDR defects. They illustrate why patch management, prevention, endpoint visibility, and investigation workflows must work as a system.

Build a measurable detection program

  1. Inventory coverage. Track protected endpoints, agent health, OS versions, sensor status, and meaningful exclusions.
  2. Model the behavior. Map priority threats to ATT&CK techniques, but write detections around observable, environment-specific hypotheses.
  3. Test safely. Exercise detections in a lab and through approved purple-team simulations. Measure telemetry arrival, rule hit, triage quality, and response time.
  4. Tune from cases. Record true positives, false positives, root causes, and tuning changes. A suppression without an expiry and owner is technical debt.
  5. Review blind spots. Reassess exclusions, new SaaS tooling, driver changes, and endpoints that fall out of policy.

Key takeaways

  • EDR is an evidence and response system, not merely an antivirus alert feed.
  • Layer prevention, atomic signals, behavioral analytics, and cross-domain hunting.
  • Engineer telemetry around investigative questions, agent health, and retention constraints.
  • Test containment actions and use automation only at an evidence threshold proportionate to business impact.
#edr#detection-engineering#windows-defense#threat-hunting
← All articlesImprove your detection coverage →