Incident Response Fundamentals
Incident Response (IR) is the disciplined, repeatable process a blue team uses to detect, analyze, contain, and recover from security incidents while preserving evidence and learning from each event. Good IR is less about heroics during a crisis and more about the preparation, tooling, and runbooks you put in place beforehand.
A critical distinction underpins everything: an event is any observable occurrence on a system or network (a login, a file write, a firewall deny). An incident is an event, or series of events, that violates or threatens security policy — for example the confirmed exfiltration of data or a successful malware detonation. Most events are benign noise; triage exists to promote the small fraction that are truly incidents.
The IR Lifecycle
Two frameworks dominate practice. NIST SP 800-61 (Rev. 2, with Rev. 3 aligning to the CSF 2.0 functions in 2025) describes four phases, while the SANS PICERL model breaks the work into six. They map cleanly onto each other; teams often speak PICERL day-to-day and cite NIST in policy documents.
| NIST SP 800-61 | SANS PICERL | Goal |
|---|---|---|
| Preparation | Preparation | Be ready: plans, tooling, logging, training |
| Detection & Analysis | Identification | Confirm an incident and scope it |
| Containment, Eradication & Recovery | Containment | Stop the bleeding without destroying evidence |
| Containment, Eradication & Recovery | Eradication | Remove the threat and its footholds |
| Containment, Eradication & Recovery | Recovery | Restore to trusted operation and monitor |
| Post-Incident Activity | Lessons Learned | Improve so it doesn't recur |
Preparation
Preparation is the phase that determines whether every other phase succeeds. You cannot analyze logs you never collected, nor isolate a host you cannot reach. Key deliverables:
- IR plan & policy — roles, authority to disconnect systems, escalation paths, and severity definitions approved by leadership.
- Runbooks / playbooks — step-by-step guides for common incident types so responders don't improvise under pressure.
- On-call & contacts — a rotation, an out-of-band comms channel (not the potentially compromised email/Slack), and legal/PR/exec contacts.
- Asset & identity inventory — you must know what you own, who owns it, and what "normal" looks like.
- Logging & visibility — centralized logs (SIEM), endpoint telemetry (EDR), network sensors, and adequate retention before the incident.
Detection & Triage
Detections originate from the SIEM (correlated log alerts), EDR (endpoint behavior), IDS/IPS (network signatures/anomalies), threat intel matches, or a human report. Each alert must be triaged. Triage answers a small set of blunt questions:
- What happened, and is it a true positive?
- When did it start (first observed activity)?
- Scope — how many hosts, accounts, and data stores are involved?
- Severity — business impact and urgency.
Severity Classification
| Severity | Example | Response |
|---|---|---|
| SEV-1 Critical | Active ransomware, confirmed data breach, domain compromise | All-hands, exec + legal notified immediately |
| SEV-2 High | Single host malware with C2, privileged account misuse | IR team engaged, containment within hours |
| SEV-3 Medium | Contained commodity malware, isolated phishing click | Standard workflow, business hours |
| SEV-4 Low | Policy violation, blocked scan, informational | Log, monitor, no urgent action |
Analysis
Analysis turns raw alerts into a coherent story. The core techniques are log analysis (pivoting across authentication, process, network, and DNS logs), timeline building (ordering events by timestamp to reconstruct the kill chain), and MITRE ATT&CK mapping (labeling observed behavior with tactics and techniques such as T1566 Phishing or T1059 Command and Scripting Interpreter).
ATT&CK mapping is powerful because it turns a messy incident into a structured hypothesis: if you saw Initial Access and Execution, you should actively hunt for Persistence and Lateral Movement rather than assuming the attacker stopped.
Analyst tip
Always normalize timestamps to UTC when building a timeline. Mixing local time zones from different log sources is one of the most common ways to draw the wrong causal conclusion during an investigation.
Containment
Containment limits damage while preserving your ability to investigate. Distinguish two horizons:
- Short-term — fast, reversible actions: network-isolate the host via EDR, block a C2 domain/IP, disable a compromised account, revoke a session token.
- Long-term — durable fixes applied while you rebuild: network segmentation, temporary firewall rules, tightened access policies.
A common containment step is to isolate the host (EDR network quarantine keeps it reachable for forensics) and disable the account rather than delete it. Crucially, preserve evidence first.
Do not just pull the plug
If you need volatile memory (running malware, injected code, encryption keys, active network connections), do NOT power off the machine. A hard shutdown destroys RAM. Capture a memory image first, then isolate. Pulling the plug is only appropriate when active destruction (e.g. ongoing wiping) outweighs the forensic loss.
Eradication
Eradication removes the adversary's presence: delete malware and dropped tools, remove persistence mechanisms (scheduled tasks, services, run keys, cron jobs, rogue accounts), patch the vulnerability that enabled initial access, and rotate all credentials that may have been exposed — including service accounts, API keys, and, after domain compromise, the krbtgt account twice.
Recovery
Recovery returns systems to trusted production. Rebuild or restore from a known-clean backup taken before the compromise, validate integrity, and phase systems back with heightened monitoring. Watch specifically for reinfection and attacker re-entry using credentials or backdoors you may have missed — attackers frequently return within days.
Forensics & Evidence Handling
Collect evidence in the order of volatility — most ephemeral first — so you don't lose data by acting slowly:
- CPU registers, cache
- Routing tables, ARP cache, process table, kernel stats, memory (RAM)
- Temporary file systems and swap
- Disk (persistent storage)
- Remote logging and monitoring data
- Physical config and network topology
- Archival media / backups
Maintain a chain of custody: document who collected each item, when, from where, and every subsequent handoff. Perform disk and memory imaging with write-blockers and work only from copies. Prove integrity by hashing evidence at acquisition and re-verifying later:
# Hash a disk image / evidence file for integrity
sha256sum evidence.dd > evidence.dd.sha256
# Later, verify nothing changed
sha256sum -c evidence.dd.sha256
# Memory acquisition (Linux, e.g. AVML) then hash
avml memory.lime && sha256sum memory.lime
Live-Response Triage Commands
Quick commands to characterize a suspect host. On Linux:
# Running processes and their trees
ps auxf
# Listening ports and active connections (modern ss, or legacy netstat)
ss -tunap
netstat -tunap
# Recent logins and reboots
last -F
lastb # failed logins
# System journal since a time window
journalctl --since "2026-09-12 08:00" --until "2026-09-12 10:00"
# Auth failures on Debian/Ubuntu
grep -i "failed\|accepted" /var/log/auth.log
# Persistence hotspots
crontab -l; ls -la /etc/cron.*; systemctl list-unit-files --state=enabled
Windows equivalents: Get-Process / tasklist for processes, Get-NetTCPConnection / netstat -ano for connections, Get-WinEvent for logs (Event ID 4624/4625 logon success/fail, 4688 process creation), and schtasks plus Autoruns for persistence.
Indicators of Compromise (IOCs)
| IOC Type | Example | Where to hunt |
|---|---|---|
| File hash | SHA-256 of a dropper | EDR, VirusTotal, endpoints |
| IP / domain | C2 server, phishing host | Firewall, proxy, DNS logs |
| URL | Malware download link | Proxy, email gateway |
| Registry / file path | Run key, dropped binary | EDR, host forensics |
| TTP (ATT&CK) | T1053 Scheduled Task | Behavioral analytics, threat hunting |
Atomic IOCs (hashes, IPs) are cheap for attackers to change; behavioral indicators and TTPs sit higher on the "Pyramid of Pain" and yield more durable detections.
Post-Incident Activity
Within days of closure, run a blameless postmortem focused on systemic causes, not individuals. Capture a factual timeline, root cause, what worked, what didn't, and concrete action items with owners. Track program health with metrics:
| Metric | Meaning |
|---|---|
| MTTD | Mean Time To Detect — dwell time before you noticed |
| MTTA | Mean Time To Acknowledge an alert |
| MTTR | Mean Time To Respond / Recover |
Common Playbooks
| Incident | Key steps |
|---|---|
| Phishing | Preserve email + headers, identify clickers, pull the message org-wide, reset exposed creds, block sender/URL |
| Ransomware | Isolate fast, identify variant, preserve a sample, find patient zero, restore from clean backups, do not rush to pay |
| Account compromise | Disable/lock, revoke sessions & tokens, reset MFA, review mailbox rules and OAuth grants, check lateral movement |
| Data exfiltration | Determine what/how much left, block channel, quantify records, preserve logs, trigger legal/breach assessment |
Communication & Legal
Decide early who to notify: internal leadership, legal, PR, affected users, regulators, cyber insurance, and law enforcement where appropriate. Regulatory clocks are strict — under the EU GDPR, controllers must notify the supervisory authority of a qualifying personal-data breach within 72 hours of becoming aware. Other regimes have their own deadlines (e.g. many U.S. states, sector rules like HIPAA, and financial disclosure rules such as the SEC's material-incident timeline). Coordinate all external messaging through legal; keep an out-of-band channel so responders aren't discussing the incident on potentially compromised systems.
Practice Exercises
- Write a one-page IR runbook for "compromised employee laptop," including containment, evidence, and eradication steps mapped to the PICERL phases.
- Load the Splunk BOTS (Boss of the SOC) or an ELK sample dataset and reconstruct an attack timeline in UTC using authentication and process logs.
- Triage a simulated phishing email: extract and interpret the headers, identify the sender infrastructure, and list the IOCs you would block.
- In a lab VM, capture a memory image (AVML/LiME) and a disk image (dd/FTK Imager), then hash both with
sha256sumand record a chain-of-custody form. - Take one investigated intrusion and map every observed behavior to MITRE ATT&CK tactics and technique IDs, then propose a detection for each.
- Run a 45-minute tabletop exercise for a ransomware scenario with defined roles, injecting new information every 10 minutes, and capture decisions and gaps in a blameless after-action report.