Incident Response Planning: A Template That Works
Incident Response Planning: A Template That Works
Change Healthcare had an incident response plan. It was filed, reviewed, and probably tested in some form of tabletop exercise. On February 21, 2024, when ALPHV/BlackCat ransomware encrypted systems processing one-third of all American healthcare claims, the plan met reality -- and response took months, not the days or hours the plan presumably envisioned. The resulting $2.4 billion in direct losses, the months-long disruption to pharmacies and providers nationwide, and the subsequent congressional hearings all point to one conclusion: having a plan is not the same as having a plan that works under pressure.
The gap between a plan that satisfies a compliance checkbox and a plan that actually reduces breach impact is not about document length or format. It is about whether the plan has been tested under conditions that approximate real stress, whether the people named in it know their roles without consulting the document, and whether the organization updates the plan based on what it learns.
This article provides a complete IR plan structure based on NIST SP 800-61 Rev 2 that is designed to function under pressure -- with specific attention to the failure modes that turned Change Healthcare's theoretical preparedness into actual catastrophe.
Why Now
Three developments have made IR plan adequacy a measurable, enforceable obligation rather than a best-practice recommendation:
- SEC Cybersecurity Rule (Dec 2023): Public companies must disclose material incidents within four business days of materiality determination. This means your IR plan must include a process for materiality assessment that can execute in hours, not weeks.
- Change Healthcare ($2.4B, Feb 2024): Demonstrated that healthcare organizations with filed IR plans can still suffer months-long response timelines when the plan has not been stress-tested against scenarios at the actual scale of the organization's interconnections.
- Federal response mandates: CISA binding operational directives and incident reporting requirements impose specific detection and response timelines on federal agencies, making IR plan testing and measurement an audit finding rather than a recommendation.
Organizations that have not pressure-tested their IR plan against a realistic scenario in the past 12 months are operating on assumption, not evidence, that they can respond within the timelines regulators now require.
IR Plan Structure
Roles and Responsibilities
Define these roles before an incident occurs. During a real incident, there is no time to negotiate who does what.
| Role | Responsibility | Backup |
|---|---|---|
| Incident Commander | Overall coordination and decision authority | Named individual + backup |
| Technical Lead | Directs investigation and containment | Senior engineer with IR experience |
| Communications Lead | Internal and external communications | Must understand legal constraints |
| Legal Counsel | Regulatory reporting, privilege, law enforcement | External counsel pre-engaged |
| Executive Sponsor | Resource authorization and business decisions | Must have actual authority |
| Scribe | Documents timeline, actions, and decisions | Dedicated -- cannot be dual-hatted |
Critical point: the scribe role is non-negotiable. During a real incident, team memory becomes unreliable within hours. The contemporaneous record is simultaneously your regulatory defense, lessons-learned input, and forensic timeline.
Incident Classification
Define severity levels with clear, objective criteria:
| Severity | Criteria | Response SLA |
|---|---|---|
| Critical | Active exfiltration, encryption in progress, safety system compromise | All hands within 30 min, executive notification immediate |
| High | Confirmed unauthorized access to sensitive data, active lateral movement | IR team assembled within 1 hour |
| Medium | Suspicious activity confirmed malicious, contained to non-critical systems | IR team notified, investigation within 4 hours |
| Low | Failed attack, isolated anomaly, contained policy violation | Documented, investigated within 24 hours |
The classification must be objective enough that an on-call engineer at 2 AM can make the severity determination without calling a meeting. If classification requires committee approval, your response time includes committee assembly time -- and attackers do not pause for quorum.
The Six Phases (Expanded from NIST SP 800-61 Rev 2's Four-Phase Lifecycle)
NIST SP 800-61 Rev 2 defines four phases -- Preparation; Detection and Analysis; Containment, Eradication, and Recovery; and Post-Incident Activity. Splitting the combined third phase into its three components yields the six working phases below.
Phase 1: Preparation
Preparation is the only phase you can execute before an incident occurs. It determines everything that follows.
- Maintain and update the IR plan -- not annually, but after every exercise and every real incident
- Train the IR team through tabletop and functional exercises (quarterly minimum)
- Ensure tools are ready: forensic workstation imaged and updated, out-of-band communication channels tested, contact lists verified (people change roles)
- Pre-engage external resources: legal counsel, forensic firm, crisis communications, law enforcement contacts, cyber insurance carrier notification procedures
- Maintain offline copies of the IR plan, contact lists, and network diagrams. If ransomware encrypts your file shares, can you still access the plan?
Phase 2: Detection and Analysis
- Identify the incident through SIEM alerts, EDR detections, user reports, or threat hunting
- Confirm: is this real or a false positive? Base decision on evidence, not hope.
- Determine scope: what systems, data, and users are affected or potentially affected?
- Classify severity using the defined criteria
- Preserve evidence: forensic images, log snapshots, memory captures before containment actions alter the state
- Begin the notification timeline assessment -- clocks may already be running
Phase 3: Containment
Two phases of containment operate on different timescales:
Short-term containment (minutes to hours):
- Isolate affected systems (network segmentation, EDR isolation, firewall rules)
- Block known attacker infrastructure (IPs, domains, C2 channels)
- Disable compromised accounts (without alerting the attacker if possible)
- Preserve evidence before containment actions overwrite volatile data
Long-term containment (hours to days):
- Apply emergency patches to exploited vulnerabilities
- Harden configurations that enabled the attack
- Implement enhanced monitoring on all systems that may have been accessed
- Build clean systems for recovery parallel to containment
Document every containment action with timestamp, actor, and rationale. This record feeds the post-incident review, regulatory inquiries, and potential litigation.
Phase 4: Eradication
- Remove all attacker persistence mechanisms (backdoors, new accounts, scheduled tasks, modified configurations)
- Identify and remediate the root cause (the vulnerability or misconfiguration that enabled initial access)
- Verify eradication through scanning, hunting, and monitoring -- absence of alerts is not evidence of absence
- Reset all potentially compromised credentials, including service accounts and API keys
The most common eradication failure: incomplete credential reset. If the attacker harvested credentials during the compromise and you reset only the accounts you know were used, they retain access through the ones you missed.
Phase 5: Recovery
- Restore affected systems from verified-clean backups (not backups that may contain attacker persistence)
- Rebuild systems from known-good images where backup integrity is uncertain
- Verify system integrity before returning to production
- Implement enhanced monitoring for recurrence (the attacker may try the same vector or use persistence you missed)
- Restore services in priority order based on business impact
Recovery is not "return to previous state." If the previous state was vulnerable, returning to it invites re-compromise. Recovery means returning to a hardened state that addresses the root cause.
Phase 6: Post-Incident Activity
- Conduct a lessons-learned review within two weeks while memory is fresh
- Identify what worked, what failed, and what was missing from the plan
- Update the IR plan based on findings -- specific, assigned changes with deadlines
- Create new detection rules to catch similar incidents and variants
- Document the full incident timeline for compliance records and audit evidence
- Assess notification obligations and execute if applicable
Communication Plan
Internal Communication
- IR team: Secure, out-of-band channel. Assume primary communication systems may be compromised. Pre-establish: Signal group, dedicated phone bridge, or physically isolated Slack workspace.
- Executive leadership: Status updates at defined intervals based on severity (Critical = hourly, High = every 4 hours). Format: one paragraph situation summary, one paragraph actions taken, one paragraph decisions needed.
- Broader organization: Need-to-know updates about service impact and required user actions (password reset, stop using a system, report suspicious activity).
- Board of directors: For material incidents, notification within hours. The SEC 4-day clock starts at materiality determination; the board should not learn about a material incident from the 8-K filing.
External Communication
- Regulators: DFARS 7012 (72 hours), HIPAA (60 days), state breach laws (30-60 days), SEC (4 business days from materiality). Map your organization to all applicable deadlines before an incident.
- Customers/partners: Per contractual obligations. Many contracts specify notification within 24-72 hours of discovery.
- Law enforcement: FBI IC3, CISA, sector-specific ISAC. Engage through legal counsel.
- Media: Through communications lead only, with legal review. No freelance statements.
Testing Your Plan
The Contrarian View on Tabletops
Quarterly tabletops are theater unless the scribe's notes change the plan within 48 hours. Most organizations run tabletop exercises to satisfy a compliance checkbox: the team discusses a scenario for two hours, someone produces a summary report, the report enters a filing system, and nothing changes. The same failure modes identified in the exercise persist until the next exercise identifies them again.
The fix is a 48-hour rule: every finding from a tabletop exercise must produce either (a) a specific plan update with an assigned owner and deadline, or (b) a documented decision that the identified risk is accepted at the executive level. If neither happens within 48 hours, the exercise was theater. Track exercise-to-change conversion rate as a metric. If it is below 50%, your exercises are producing paper, not preparedness.
Testing Levels
Tabletop exercises (quarterly): Walk through a scenario verbally with the IR team. No actual systems involved. Duration: 2-4 hours. Focus: decision-making, communication, coordination. Low cost, high insight.
Functional exercises (semi-annually): Simulate an incident using test systems. The IR team performs actual containment and investigation steps. Duration: 4-8 hours. Focus: tool proficiency, process execution, timing. Tests whether the team can actually do what the plan says they should do.
Full-scale exercises (annually): Simulate a major incident with cross-functional participation: technical team, communications, legal, executive leadership, external partners. Duration: 1-2 days. Focus: organizational coordination, escalation paths, external communication. Tests the full system, not just the security team.
Purple team operations (as scheduled): Red team executes real (controlled) attacks against production systems. Blue team detects and responds using the IR plan. Duration: 1-5 days. Focus: detection capability, response speed, plan adequacy against real techniques. The highest-fidelity test available.
Metrics That Matter
| Metric | Target | What It Tests |
|---|---|---|
| Time from detection to severity classification | Under 30 minutes | Phase 2 efficiency |
| Time from classification to team assembly | Under 1 hour (High/Critical) | Escalation paths |
| Time from assembly to initial containment | Under 4 hours | Phase 3 readiness |
| Exercise-to-plan-change conversion rate | Above 50% | Exercise value |
| Notification deadline compliance | 100% | Communication plan |
| Plan update freshness | Updated within 30 days of last exercise/incident | Maintenance discipline |
Common IR Plan Failures
The plan lives in the encrypted file share. If ransomware encrypts your systems, can you access the plan? Maintain offline copies: printed binders, USB drives in a safe. The plan must survive the incident it addresses.
Contact lists are stale. People change roles quarterly. Verify contact lists monthly -- discovering the "legal counsel" number reaches someone who left six months ago wastes critical hours.
No one has authority to make expensive decisions fast. Engaging a $50K/week forensic firm or shutting production requires authority. If that authority needs a committee that cannot assemble until Monday, your Friday-night incident stalls 48 hours.
The plan assumes prior phases worked. What happens when containment fails or the "clean" backup contains the backdoor? Build decision trees, not linear procedures.
Post-incident findings change nothing. Track which findings produced plan changes. Unactioned findings are organizational amnesia.
Key Takeaways
- An IR plan that has not been tested under stress is an assumption, not a capability -- Change Healthcare's $2.4B loss proved that filed plans and actual preparedness are different things
- The scribe role is non-negotiable: contemporaneous documentation is simultaneously your regulatory defense, forensic timeline, and lessons-learned input
- Severity classification must be objective enough for a 2 AM on-call determination without committee assembly
- Contain in two phases: short-term (stop the bleeding in minutes) and long-term (harden for recovery over hours/days)
- Quarterly tabletops are theater unless findings produce plan changes within 48 hours -- track exercise-to-change conversion rate
- The plan must survive the incident: maintain offline copies because ransomware does not spare your file shares
- Recovery means returning to a hardened state, not the vulnerable state that enabled the compromise
- The SEC's 4-day materiality rule means your plan must include a rapid materiality assessment process, not a weeks-long deliberation
Frequently Asked Questions
How often should we update the IR plan?
Update after every exercise, every real incident, every significant organizational change (new systems, new regulatory obligations, personnel changes), and at minimum annually even if none of those triggers occur. The test for staleness: pick any page of the plan. Are the contact names current? Are the systems listed still in production? Are the procedures still executable with current tools? If any answer is no, the plan is stale. Track "last updated" date as an auditable metric -- most frameworks require documented review cadence.
What is the minimum team size for effective incident response?
You need at least four distinct roles filled (incident commander, technical lead, communications/legal, scribe) even if those roles are filled by two or three people in a small organization. What you cannot do: have the technical lead also serve as scribe (they are too busy investigating to document), or have the incident commander also handle communications (conflicting priorities during triage). For organizations under 500 employees, cross-trained personnel filling two non-conflicting roles is acceptable. For organizations with compliance obligations (HIPAA, DFARS, FedRAMP), your assessor will verify that named individuals can actually perform their assigned roles.
How do we handle an incident that spans multiple cloud providers and on-premises systems?
Pre-build provider-specific runbooks. For each environment: (1) break-glass IAM credentials tested quarterly, (2) documented procedures for log export, snapshot creation, and network isolation, (3) pre-authorized isolation rules that require no approval delay, (4) shared responsibility boundary documentation. The #1 multi-cloud IR failure: discovering during the incident that you lack permissions to collect evidence from a specific environment.
What should the IR plan say about paying ransomware demands?
Document a decision framework, not a blanket policy. Factors: (1) whether clean backups exist, (2) whether payment is legal (OFAC sanctions), (3) whether the actor provides working decryptors, (4) business impact of extended downtime vs. cost, (5) insurance coverage. Pre-define decision authority (typically CEO with board notification). The FBI recommends against payment but does not prohibit it. OFAC sanctions may prohibit payment to specific groups regardless of impact.
How do we demonstrate IR plan adequacy to auditors and assessors?
Auditors evaluate three things: (1) the plan document itself (roles, phases, procedures, communication plan), (2) evidence of testing (exercise reports with dates, participants, findings, and resulting plan changes), (3) evidence of maintenance (version history, update logs, contact list verification records). The strongest evidence: a post-incident report showing the plan was executed during a real event with documented timeline, decisions, and lessons learned that produced specific plan updates. The weakest evidence: an annual plan review with no test documentation and no changes year-over-year. Assessors for FedRAMP IR-8 and CMMC explicitly require exercise evidence, not just the plan document.
How Advisedly Helps
Advisedly includes incident response workflow tools that guide your team through the NIST 800-61 phases with structured checklists, automated evidence collection, timeline documentation, and regulatory deadline tracking across HIPAA, DFARS, SEC, and state breach laws simultaneously. The platform tracks exercise completion and plan update history as compliance evidence for SOC 2, FedRAMP, and CMMC, closing the gap between "we have a plan" and "we can prove the plan works." Contact begin@advisedly.ai to see how IR workflow automation works in practice.
<!-- LI hook: Change Healthcare had an IR plan. Response still took months and cost $2.4 billion. -->