contentintech
Learn/cybersecurity/Incident Response
Intermediate~22 min read

Incident Response

The incident response lifecycle, detection and triage, containment, forensic evidence handling, and post-incident review.

IR LifecycleForensicsDetectionPlaybooks

Incident Response Fundamentals

Incident Response (IR) is the disciplined, repeatable process a blue team uses to detect, analyze, contain, and recover from security incidents while preserving evidence and learning from each event. Good IR is less about heroics during a crisis and more about the preparation, tooling, and runbooks you put in place beforehand.

A critical distinction underpins everything: an event is any observable occurrence on a system or network (a login, a file write, a firewall deny). An incident is an event, or series of events, that violates or threatens security policy — for example the confirmed exfiltration of data or a successful malware detonation. Most events are benign noise; triage exists to promote the small fraction that are truly incidents.

The IR Lifecycle

Two frameworks dominate practice. NIST SP 800-61 (Rev. 2, with Rev. 3 aligning to the CSF 2.0 functions in 2025) describes four phases, while the SANS PICERL model breaks the work into six. They map cleanly onto each other; teams often speak PICERL day-to-day and cite NIST in policy documents.

NIST SP 800-61 SANS PICERL Goal
PreparationPreparationBe ready: plans, tooling, logging, training
Detection & AnalysisIdentificationConfirm an incident and scope it
Containment, Eradication & RecoveryContainmentStop the bleeding without destroying evidence
Containment, Eradication & RecoveryEradicationRemove the threat and its footholds
Containment, Eradication & RecoveryRecoveryRestore to trusted operation and monitor
Post-Incident ActivityLessons LearnedImprove so it doesn't recur

Preparation

Preparation is the phase that determines whether every other phase succeeds. You cannot analyze logs you never collected, nor isolate a host you cannot reach. Key deliverables:

  1. IR plan & policy — roles, authority to disconnect systems, escalation paths, and severity definitions approved by leadership.
  2. Runbooks / playbooks — step-by-step guides for common incident types so responders don't improvise under pressure.
  3. On-call & contacts — a rotation, an out-of-band comms channel (not the potentially compromised email/Slack), and legal/PR/exec contacts.
  4. Asset & identity inventory — you must know what you own, who owns it, and what "normal" looks like.
  5. Logging & visibility — centralized logs (SIEM), endpoint telemetry (EDR), network sensors, and adequate retention before the incident.

Detection & Triage

Detections originate from the SIEM (correlated log alerts), EDR (endpoint behavior), IDS/IPS (network signatures/anomalies), threat intel matches, or a human report. Each alert must be triaged. Triage answers a small set of blunt questions:

  1. What happened, and is it a true positive?
  2. When did it start (first observed activity)?
  3. Scope — how many hosts, accounts, and data stores are involved?
  4. Severity — business impact and urgency.

Severity Classification

Severity Example Response
SEV-1 CriticalActive ransomware, confirmed data breach, domain compromiseAll-hands, exec + legal notified immediately
SEV-2 HighSingle host malware with C2, privileged account misuseIR team engaged, containment within hours
SEV-3 MediumContained commodity malware, isolated phishing clickStandard workflow, business hours
SEV-4 LowPolicy violation, blocked scan, informationalLog, monitor, no urgent action

Analysis

Analysis turns raw alerts into a coherent story. The core techniques are log analysis (pivoting across authentication, process, network, and DNS logs), timeline building (ordering events by timestamp to reconstruct the kill chain), and MITRE ATT&CK mapping (labeling observed behavior with tactics and techniques such as T1566 Phishing or T1059 Command and Scripting Interpreter).

ATT&CK mapping is powerful because it turns a messy incident into a structured hypothesis: if you saw Initial Access and Execution, you should actively hunt for Persistence and Lateral Movement rather than assuming the attacker stopped.

Analyst tip

Always normalize timestamps to UTC when building a timeline. Mixing local time zones from different log sources is one of the most common ways to draw the wrong causal conclusion during an investigation.

Containment

Containment limits damage while preserving your ability to investigate. Distinguish two horizons:

  1. Short-term — fast, reversible actions: network-isolate the host via EDR, block a C2 domain/IP, disable a compromised account, revoke a session token.
  2. Long-term — durable fixes applied while you rebuild: network segmentation, temporary firewall rules, tightened access policies.

A common containment step is to isolate the host (EDR network quarantine keeps it reachable for forensics) and disable the account rather than delete it. Crucially, preserve evidence first.

Do not just pull the plug

If you need volatile memory (running malware, injected code, encryption keys, active network connections), do NOT power off the machine. A hard shutdown destroys RAM. Capture a memory image first, then isolate. Pulling the plug is only appropriate when active destruction (e.g. ongoing wiping) outweighs the forensic loss.

Eradication

Eradication removes the adversary's presence: delete malware and dropped tools, remove persistence mechanisms (scheduled tasks, services, run keys, cron jobs, rogue accounts), patch the vulnerability that enabled initial access, and rotate all credentials that may have been exposed — including service accounts, API keys, and, after domain compromise, the krbtgt account twice.

Recovery

Recovery returns systems to trusted production. Rebuild or restore from a known-clean backup taken before the compromise, validate integrity, and phase systems back with heightened monitoring. Watch specifically for reinfection and attacker re-entry using credentials or backdoors you may have missed — attackers frequently return within days.

Forensics & Evidence Handling

Collect evidence in the order of volatility — most ephemeral first — so you don't lose data by acting slowly:

  1. CPU registers, cache
  2. Routing tables, ARP cache, process table, kernel stats, memory (RAM)
  3. Temporary file systems and swap
  4. Disk (persistent storage)
  5. Remote logging and monitoring data
  6. Physical config and network topology
  7. Archival media / backups

Maintain a chain of custody: document who collected each item, when, from where, and every subsequent handoff. Perform disk and memory imaging with write-blockers and work only from copies. Prove integrity by hashing evidence at acquisition and re-verifying later:

# Hash a disk image / evidence file for integrity
sha256sum evidence.dd > evidence.dd.sha256

# Later, verify nothing changed
sha256sum -c evidence.dd.sha256

# Memory acquisition (Linux, e.g. AVML) then hash
avml memory.lime && sha256sum memory.lime

Live-Response Triage Commands

Quick commands to characterize a suspect host. On Linux:

# Running processes and their trees
ps auxf

# Listening ports and active connections (modern ss, or legacy netstat)
ss -tunap
netstat -tunap

# Recent logins and reboots
last -F
lastb            # failed logins

# System journal since a time window
journalctl --since "2026-09-12 08:00" --until "2026-09-12 10:00"

# Auth failures on Debian/Ubuntu
grep -i "failed\|accepted" /var/log/auth.log

# Persistence hotspots
crontab -l; ls -la /etc/cron.*; systemctl list-unit-files --state=enabled

Windows equivalents: Get-Process / tasklist for processes, Get-NetTCPConnection / netstat -ano for connections, Get-WinEvent for logs (Event ID 4624/4625 logon success/fail, 4688 process creation), and schtasks plus Autoruns for persistence.

Indicators of Compromise (IOCs)

IOC Type Example Where to hunt
File hashSHA-256 of a dropperEDR, VirusTotal, endpoints
IP / domainC2 server, phishing hostFirewall, proxy, DNS logs
URLMalware download linkProxy, email gateway
Registry / file pathRun key, dropped binaryEDR, host forensics
TTP (ATT&CK)T1053 Scheduled TaskBehavioral analytics, threat hunting

Atomic IOCs (hashes, IPs) are cheap for attackers to change; behavioral indicators and TTPs sit higher on the "Pyramid of Pain" and yield more durable detections.

Post-Incident Activity

Within days of closure, run a blameless postmortem focused on systemic causes, not individuals. Capture a factual timeline, root cause, what worked, what didn't, and concrete action items with owners. Track program health with metrics:

Metric Meaning
MTTDMean Time To Detect — dwell time before you noticed
MTTAMean Time To Acknowledge an alert
MTTRMean Time To Respond / Recover

Common Playbooks

Incident Key steps
PhishingPreserve email + headers, identify clickers, pull the message org-wide, reset exposed creds, block sender/URL
RansomwareIsolate fast, identify variant, preserve a sample, find patient zero, restore from clean backups, do not rush to pay
Account compromiseDisable/lock, revoke sessions & tokens, reset MFA, review mailbox rules and OAuth grants, check lateral movement
Data exfiltrationDetermine what/how much left, block channel, quantify records, preserve logs, trigger legal/breach assessment

Communication & Legal

Decide early who to notify: internal leadership, legal, PR, affected users, regulators, cyber insurance, and law enforcement where appropriate. Regulatory clocks are strict — under the EU GDPR, controllers must notify the supervisory authority of a qualifying personal-data breach within 72 hours of becoming aware. Other regimes have their own deadlines (e.g. many U.S. states, sector rules like HIPAA, and financial disclosure rules such as the SEC's material-incident timeline). Coordinate all external messaging through legal; keep an out-of-band channel so responders aren't discussing the incident on potentially compromised systems.

Practice Exercises

  1. Write a one-page IR runbook for "compromised employee laptop," including containment, evidence, and eradication steps mapped to the PICERL phases.
  2. Load the Splunk BOTS (Boss of the SOC) or an ELK sample dataset and reconstruct an attack timeline in UTC using authentication and process logs.
  3. Triage a simulated phishing email: extract and interpret the headers, identify the sender infrastructure, and list the IOCs you would block.
  4. In a lab VM, capture a memory image (AVML/LiME) and a disk image (dd/FTK Imager), then hash both with sha256sum and record a chain-of-custody form.
  5. Take one investigated intrusion and map every observed behavior to MITRE ATT&CK tactics and technique IDs, then propose a detection for each.
  6. Run a 45-minute tabletop exercise for a ransomware scenario with defined roles, injecting new information every 10 minutes, and capture decisions and gaps in a blameless after-action report.

Section navigation